AI Detectors Fail the Students Who Follow the Rules

I’ve said for years that AI detectors are a poor foundation for academic integrity. A new study gives me the hardest evidence yet, and it goes further than I expected. Sun, Liao, and Ma (2026), writing in Computers & Education, ran one of the largest detector evaluations to date, and the picture they paint should worry anyone still treating a detection report as proof of cheating.

The scale alone earns attention. The team built three datasets totalling over 280,000 paired samples, with the human-written half drawn from student work archived between 2016 and 2021, before ChatGPT existed. They then tested 13 detectors, both commercial and open-source, across coursework, theses, and engineering code.

Detection performance swung wildly by genre. The tools did passably on long theses, worse on short coursework, and barely above a coin flip on code. GPTZero, a market leader, missed roughly three-quarters of AI-generated code. Several detectors scored below random, which means a coin toss would serve a teacher better.

The robustness test is where the case collapses entirely. Sun et al. interviewed 25 students about how they actually lower their AI scores, then tested those real-world edits. Simple synonym swaps pushed over half of AI text past the detectors. A combined editing strategy drove the evasion rate to 88%. Nearly nine in ten polished AI assignments sailed through undetected, and the supposedly strongest tool failed hardest.

This tracks with everything I’ve argued about the detection arms race. You cannot win it. Students hold the cheaper move.

AI Detectors

The Part That Crosses Into Injustice

Here is the finding that moved this paper from “detectors are unreliable” to “detectors are unfair.” The authors found a systematic bias against STEM writing. Technical prose is formulaic by design, full of passive voice, fixed phrasing, and standardized structure, and that formulaic quality reads as machine-like to a statistical detector. So the students most likely to be falsely flagged are the ones following disciplinary conventions most faithfully.

The misclassification of rigorous human work, the authors write, “is not only unfair but also a potential penalty for upholding academic standards” (p. 17). In other words, a diligent engineering student who writes clean, conventional code or a precise methods section carries a higher false-accusation risk than a humanities peer writing with flair.

This compounds a bias the field already knew about. The authors build on Liang and colleagues’ Stanford work, which found non-native English speakers flagged at rates as high as 61%, since their lower lexical variety reads as low perplexity. The new paper extends that penalty from individuals to whole disciplines. Two groups of blameless students, both punished for how they write, not for what they did.

The Incentives Run Backwards

There’s a structural twist that deserves more attention than it will get. Detection gets harder as tasks get harder. Easy factual prompts produce predictable AI text that’s simple to catch. Complex reasoning and long codebases pull AI output close to expert human work, so it slips through.

The consequence is a reverse-selection effect. Students face the lowest detection risk on exactly the high-stakes, complex work that matters most, and the highest risk on trivial formative tasks. The system punishes small genuine slips and waves through ambitious concealment. No assessment regime should be built on an incentive structure that runs that far backwards.

One more result complicates the usual assumptions. The expensive commercial tools were not reliably better than free open-source ones. Several paid detectors posted false-positive rates above 40% on short tasks, while an open-source zero-shot model proved the most stable.

The marketing claims of 99% accuracy do not survive contact with real student writing. Sun et al. align their results with the multi-centre evaluation by Weber-Wulff et al. (2023), which concluded that “no universal detection solution adaptable to all academic genres currently exists” (p. 16). That line, which the authors quote from another study, has aged into consensus.

How I Read This in 2026

The authors are clear that the fix is pedagogical, not technical. Their core argument is that “the future of higher education lies not in building higher technical walls, but in redefining the essence of education” (p. 19). I’ve been making a version of that case on this blog for two years.

The practical move is to demote the detector. A detection score should be one weak signal among many, never the smoking gun. Sun et al. argue for triangulating it with version histories, drafts, and oral defences, which connects directly to the assessment-validity work from Corbin et al. (2025a/2025b/2025c/2025d), Dawson et al. (2024) and Bearman et al. (20232024) that anchors much of my thinking. When detection can’t anchor integrity, the validity of the task itself has to carry the weight. The “assessment twin” idea from Roe and colleagues (2026) points the same direction, toward process evidence over a forensic verdict on the final text.

This is the pedagogy point I always land on, sharpened by hard data. The tool cannot tell you whether learning happened. The task design can. Build assessments around visible reasoning, defended choices, and the development trajectory, and the detector shrinks to a minor diagnostic, no longer a courtroom.

References

  • Bearman, M., Nieminen, J. H., & Ajjawi, R. (2023). Designing assessment in a digital world: An organising framework. _Assessment & Evaluation in Higher Education_, 48(3), 291-304. https://doi.org/10.1080/02602938.2022.2069674 .
  • Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. _Assessment & Evaluation in Higher Education, 49_(6), 893-905. https://doi.org/10.1080/02602938.2024.2335321
  • Corbin, T., Bearman, M., Boud, D., & Dawson, P. (2025a). The wicked problem of AI and assessment. Assessment & Evaluation in Higher Education. 1–17. https://doi.org/10.1080/02602938.2025.2553340
  • Corbin, T., Dawson, P., Nicola-Richmond, K., & Partridge, H. (2025b). ‘Where’s the line? It’s an absurd line’: Towards a framework for acceptable uses of AI in assessment. Assessment & Evaluation in Higher Education, 50(5), 705-717. https://doi.org/10.1080/02602938.2025.2456207
  • Corbin, T., Dawson, P., & Liu, D. (2025c). Talk is cheap: Why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. https://doi.org/10.1080/02602938.2025.2503964
  • Corbin, T., Tai, J., & Flenady, G. (2025d). Understanding the place and value of GenAI feedback: A recognition-based framework. Assessment & Evaluation in Higher Education, 50(5), 718–731. https://doi.org/10.1080/02602938.2025.2459641
  • Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. _Assessment & Evaluation in Higher Education_, 49(7), 1005–1016. https://doi.org/10.1080/02602938.2024.2386662 
  • Roe, J., Perkins, M., & Giray, L. (2025). Assessment twins: A protocol for AI-vulnerable summative assessment. arXiv preprint arXiv:2510.02929.
  • Sun, Y., Liao, Y., & Ma, X. (2026). Trusting AI to detect AI? A systematic evaluation of the reliability and robustness of current AIGC detection tools for student academic work. Computers & Education, 249, Article 105616. https://doi.org/10.1016/j.compedu.2026.105616

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top