DigitalDUAL-USE

Detector Overtrust

What it is

Treating the output of an AI-content detector as proof that a text or image is or is not machine-made, whether to accuse a person, to clear a fake, or to discredit authentic material.

How it works

Detectors return a confident-looking percentage, and people read a number from software as a measurement. It is a statistical guess. Text detectors largely score how predictable the wording is, so plain, formulaic, or second-language prose looks machine-like. Liang and colleagues (2023) found that seven detectors flagged more than half of TOEFL essays by non-native English writers as AI-generated while scoring U.S. eighth-graders' essays almost perfectly, and that light rewording defeated the detectors. OpenAI withdrew its own classifier on July 20, 2023, citing a low rate of accuracy; it had caught only 26 percent of AI text and mislabeled 9 percent of human text. Weber-Wulff and colleagues (2023) tested fourteen tools and found none reliable. The manipulation runs in both directions: a false positive can be used to accuse an innocent writer, and a false negative or a cherry-picked result can be used to launder a fake or to call real evidence synthetic.

Real-world examples

  • In May 2023 an instructor at Texas A&M University-Commerce pasted students' essays into ChatGPT, asked whether it had written them, and temporarily withheld grades for a class on the basis of its answers; the university said no student ultimately failed or was barred from graduating over it.
  • In August 2023 Vanderbilt University announced it had disabled Turnitin's AI detection feature, citing concerns about false positives and the lack of transparency about how the tool worked.
  • OpenAI's AI Text Classifier, launched January 31, 2023, was taken offline on July 20, 2023 with a note citing its low rate of accuracy.
  • During the Israel-Hamas war in October 2023, a free online image detector labeled a photograph released by the Israeli government of a victim of the October 7 attack as AI-generated; the result was shared widely as proof of fabrication, while forensic expert Hany Farid told 404 Media the image showed no signs of AI generation.

Ethical guidelines

Where the line is

Using a detector as one weak signal inside a fuller verification process, with its error rates disclosed and a human decision at the end, is legitimate triage. It becomes manipulation or negligence when the score is presented as proof, when it is used to accuse or exonerate without other evidence, or when someone shops among tools until one returns the verdict they wanted.

  • A detector score is a lead, never a verdict. Do not penalize a student, employee, or source on a score alone.
  • Anyone deploying a detector owes the people judged by it the false-positive rate, measured on writers like them.
  • Do not publish a detector screenshot as a fact-check in either direction. Show provenance and corroboration instead.
  • Vendors should not market accuracy figures drawn from test sets that exclude non-native writers and edited text.

How to defend against it

  • If you are accused on the basis of a detector: ask for the tool's documented false-positive rate, point to Liang et al. (2023) and to OpenAI's withdrawal of its own classifier, and offer process evidence such as drafts, version history, and notes. Ask to discuss the work in person.
  • If you teach or manage: assess process rather than policing prose. Use drafts, oral follow-up, and in-class writing; treat a detector flag as a reason for a conversation, not a finding.
  • If you are verifying media: use detectors last and only alongside provenance (who first posted it, when, from where), reverse image search, and corroborating footage. Seek a named forensic analyst for anything consequential.
  • When someone shares a detector result to prove a point, ask which tool, what confidence, and whether a second method agrees. Ask who benefits from the verdict.
  • Remember the asymmetry: a clean result does not prove authenticity, and a flagged result does not prove fakery.

References

  1. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779 · link
    More than half of non-native TOEFL essays misclassified as AI-generated; simple rewording defeats detection.
  2. OpenAI (2023). New AI classifier for indicating AI-written text. OpenAI, January 31, 2023, updated July 20, 2023 with the discontinuation notice · link
    The 26 percent true-positive and 9 percent false-positive figures and the withdrawal for low accuracy.
  3. 404 Media (2023). AI Images Detectors Are Being Used to Discredit the Real Horrors of War. 404 Media, October 2023 · link
    A detector false positive on an authentic war photograph and its use to discredit real evidence.
  4. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26 · link
    Evaluation of fourteen detection tools finding them neither accurate nor reliable.
Last reviewed
Suggest a correction

Detect Detector Overtrust in any text

Paste any message, email, or article into our free Manipulation Detector to see if Detector Overtrust or other techniques are being used on you.

Related Articles