Are AI detectors accurate?

Better than a coin toss, worse than their marketing. The real problem isn’t the headline accuracy – it’s what happens when you run a detector over a room full of honest people.

The base-rate problem, in one simulator

Set how many essays you mark, how many you think are AI-written and how good the detector is. Each dot is an essay.

AI, flagged (18)Human, wrongly flagged (4)AI, missed (2)Human, cleared

Of 22 flagged essays, 4 were written by students. A flag is right 82% of the time.

Move “share actually AI-written” down to 2% and watch: even a detector with a 2% false-positive rate then produces about as many wrong flags as right ones.

What the evidence says

  • OpenAI’s own classifier (January 2023) correctly labelled only 26% of AI-written text as “likely AI” and wrongly labelled 9% of human text. OpenAI withdrew it in July 2023, citing its low rate of accuracy.
  • Non-native writers are flagged far more. A Stanford study (Liang et al., 2023, published in Patterns) found seven popular detectors misclassified over half of TOEFL essays written by non-native speakers as AI-generated, while scoring US eighth-grade essays almost perfectly.
  • Paraphrasing defeats detectors. Researchers at the University of Maryland (Sadasivan et al., 2023) showed that running AI text through a paraphrasing model sharply cut detection rates across several detectors.
  • Vendors say it themselves. Turnitin states that its AI writing indicator should not be used as the sole basis for adverse action, and hides scores between 1% and 19% because of higher false-positive rates at low levels.
  • Some universities switched it off. Vanderbilt University disabled Turnitin’s AI detector in 2023, explaining that even a 1% false-positive rate would mean hundreds of wrongly flagged papers a year.

Why short texts are the worst case

Every signal is an average. Over 50 words, three unusual sentences can swing the result either way. Over 800 words, patterns stabilise. VerifyWrite shows a confidence level for exactly this reason and pulls short-text scores toward the middle.

How to use a detector responsibly

  1. Treat the score as a reason to look closer, never as evidence.
  2. Look at which sentences are flagged and why.
  3. Seek independent evidence: drafts, version history, a conversation.
  4. Be most careful with second-language writers and short answers.

See the five-step process for essays, or what Turnitin’s AI score does and doesn’t tell you.

Questions people ask

Are AI detectors accurate?

They are right more often than chance, but not reliable enough to act on alone. Accuracy drops on short texts, edited AI output, non-native English writing and newer models. Even a small false-positive rate produces many wrong flags across a large class.

What is a false positive in AI detection?

A human-written text that the detector labels as AI-generated. It is the most harmful error, because it can lead to an honest person being accused of misconduct.

Who is most likely to be falsely flagged?

Writers whose prose is predictable to a language model: non-native English speakers (Liang et al., 2023, found over half of TOEFL essays were misclassified), neurodivergent writers who favour consistent structure, and students trained on rigid essay templates.

What should I do if I’m falsely accused?

Stay calm and gather process evidence: document version history, drafts, notes, browser history for your research, and earlier writing for comparison. Ask what evidence exists beyond the detector score, and cite the vendor’s own guidance that scores should not be used alone.