Detector Overtrust
What it is
Treating the output of an AI-content detector as proof that a text or image is or is not machine-made, whether to accuse a person, to clear a fake, or to discredit authentic material.
How it works
Real-world examples
- •In May 2023 an instructor at Texas A&M University-Commerce pasted students' essays into ChatGPT, asked whether it had written them, and temporarily withheld grades for a class on the basis of its answers; the university said no student ultimately failed or was barred from graduating over it.
- •In August 2023 Vanderbilt University announced it had disabled Turnitin's AI detection feature, citing concerns about false positives and the lack of transparency about how the tool worked.
- •OpenAI's AI Text Classifier, launched January 31, 2023, was taken offline on July 20, 2023 with a note citing its low rate of accuracy.
- •During the Israel-Hamas war in October 2023, a free online image detector labeled a photograph released by the Israeli government of a victim of the October 7 attack as AI-generated; the result was shared widely as proof of fabrication, while forensic expert Hany Farid told 404 Media the image showed no signs of AI generation.
Ethical guidelines
Using a detector as one weak signal inside a fuller verification process, with its error rates disclosed and a human decision at the end, is legitimate triage. It becomes manipulation or negligence when the score is presented as proof, when it is used to accuse or exonerate without other evidence, or when someone shops among tools until one returns the verdict they wanted.
- ●A detector score is a lead, never a verdict. Do not penalize a student, employee, or source on a score alone.
- ●Anyone deploying a detector owes the people judged by it the false-positive rate, measured on writers like them.
- ●Do not publish a detector screenshot as a fact-check in either direction. Show provenance and corroboration instead.
- ●Vendors should not market accuracy figures drawn from test sets that exclude non-native writers and edited text.
How to defend against it
- ►If you are accused on the basis of a detector: ask for the tool's documented false-positive rate, point to Liang et al. (2023) and to OpenAI's withdrawal of its own classifier, and offer process evidence such as drafts, version history, and notes. Ask to discuss the work in person.
- ►If you teach or manage: assess process rather than policing prose. Use drafts, oral follow-up, and in-class writing; treat a detector flag as a reason for a conversation, not a finding.
- ►If you are verifying media: use detectors last and only alongside provenance (who first posted it, when, from where), reverse image search, and corroborating footage. Seek a named forensic analyst for anything consequential.
- ►When someone shares a detector result to prove a point, ask which tool, what confidence, and whether a second method agrees. Ask who benefits from the verdict.
- ►Remember the asymmetry: a clean result does not prove authenticity, and a flagged result does not prove fakery.
References
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779 · linkMore than half of non-native TOEFL essays misclassified as AI-generated; simple rewording defeats detection.
- OpenAI (2023). New AI classifier for indicating AI-written text. OpenAI, January 31, 2023, updated July 20, 2023 with the discontinuation notice · linkThe 26 percent true-positive and 9 percent false-positive figures and the withdrawal for low accuracy.
- 404 Media (2023). AI Images Detectors Are Being Used to Discredit the Real Horrors of War. 404 Media, October 2023 · linkA detector false positive on an authentic war photograph and its use to discredit real evidence.
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26 · linkEvaluation of fourteen detection tools finding them neither accurate nor reliable.
Related Articles
Dark Patterns in UX: How Apps Manipulate Your Behavior
Subscription traps, misleading interfaces, and engineered addiction. Understanding the persuasion techniques built into the apps you use every day.
OSINT for Beginners: Open Source Intelligence Explained
Open Source Intelligence (OSINT) uses publicly available data to gather actionable insights. Here is a beginner-friendly guide to what OSINT is and how it is used.