LumenWrite

Are AI detectors accurate?

They're a useful signal, not proof. Here is what the published evidence says, when a score deserves less trust, and how students and teachers can use one fairly.

Updated October 3, 2026

How accurate are AI detectors? The short answer

AI detectors are accurate enough to be a useful signal and not accurate enough to be proof. They do best on long, unedited AI drafts. They do worse on short texts, edited drafts, writing that mixes human and AI parts, and formal or non-native English, and they sometimes flag writing a person produced entirely alone. How far you can trust a score depends on the text and on the tool.

That is how we describe our own free AI detector too: a guide, not a verdict. We haven't published our own accuracy benchmark yet, so this page doesn't quote one. Everything below comes from published research and from what detector makers have said about their own tools. For the mechanics behind the scores, start with how AI detectors work.

False positives vs false negatives: a worked example

A detector can be wrong in two directions. A false positive is human writing flagged as AI. A false negative is AI writing that passes as human. The two pull against each other: tune a tool to catch more AI and it flags more humans; tune it to protect humans and more AI slips through.

A single percentage hides what this means for real people, so run some numbers. They are invented for illustration and don't describe any real tool. Say a school checks 1,000 essays. 100 were written with AI and 900 were not. The detector catches 90% of AI essays and wrongly flags 1% of human ones.

  • It flags 90 of the 100 AI essays and misses 10. Those 10 are false negatives.
  • It also flags 9 of the 900 human essays. Those 9 are false positives.
  • Of the 99 flagged essays, 9 belong to students who did nothing wrong: about one flag in eleven.

Now suppose only 20 of the 1,000 essays used AI. The same tool flags 18 of them, plus about 10 of the 980 human essays. More than a third of the flags are now wrong, and the detector didn't change at all. Only the mix of essays did. So "1% false positives" never means "a flag is 99% likely to be right".

Are AI detectors accurate? What the published evidence says

OpenAI withdrew its own AI Text Classifier

OpenAI released an AI Text Classifier in January 2023. On its own test set, it correctly labeled 26% of AI-written text as "likely AI-written" and labeled human writing as AI 9% of the time. On July 20, 2023, OpenAI took it offline, citing its low rate of accuracy. The company behind ChatGPT could not build a text detector it was willing to keep running.

Stanford: non-native English writers flagged far more often

A Stanford team led by Weixin Liang and James Zou tested seven widely used detectors on 91 TOEFL essays by non-native English speakers and 88 essays by US eighth graders. The detectors were near-perfect on the eighth-grade essays but labeled an average of 61% of the TOEFL essays as AI. Of the 91, 89 were flagged by at least one detector.

The likely cause is word choice: writers with a smaller English vocabulary pick more predictable words, which is exactly what detectors read as AI. When the researchers had ChatGPT enrich the vocabulary of those same essays, the average false positive rate fell from 61% to 12%. The detectors were reacting to style, not to who did the writing.

A test of 14 tools: "neither accurate nor reliable"

An international team led by Debora Weber-Wulff tested 14 detection tools, 12 publicly available ones plus Turnitin and PlagiarismCheck, on human, AI-generated, edited, paraphrased and machine-translated texts. None reached 80% accuracy, and only five passed 70%. The tools leaned toward calling text human, so they missed AI more often than they accused people. Machine-translated human writing still raised false positives, and editing or paraphrasing AI text made detection much worse. The authors concluded the tools are "neither accurate nor reliable".

What Turnitin says about its own false positive rate

Turnitin publishes its own figures. In a May 2023 update, its chief product officer said the document false positive rate is under 1% for documents where it finds more than 20% AI writing. About 4% of the individual sentences it highlights may be human-written, most of them close to real AI text. It also saw more false positives when it detected less than 20% AI, so it began marking those scores with an asterisk.

Turnitin tells instructors that because its false positive rate is not zero, they need to apply their own judgment, their knowledge of the student and the context of the assignment.

Universities that switched AI detection off

In August 2023, Vanderbilt University disabled Turnitin's AI detector. Its reasoning was arithmetic: Vanderbilt submitted 75,000 papers to Turnitin in 2022, so a 1% false positive rate could have meant around 750 papers wrongly labeled. It also cited the bias against non-native English writers and how little Turnitin had said about how the tool works.

Curtin University in Australia turned off the same feature from 1 January 2026, while keeping Turnitin's regular text-matching checks.

What affects AI detector accuracy

The published results line up with how detectors work. Trust a score less when any of these apply:

  • Length. Short texts carry little signal. OpenAI called its classifier very unreliable below 1,000 characters, and Turnitin raised its minimum from 150 to 300 words for the same reason. Our detector doesn't score anything under 40 words.
  • Editing. Each sentence a person rewrites dilutes the signal. In the 14-tool test, accuracy dropped sharply on manually edited and paraphrased AI text.
  • Language. Many detectors are built mainly for English. OpenAI said its classifier did significantly worse in other languages, and ours works best on English too. Machine-translated text adds false positives.
  • Formal or formulaic writing. Lab reports, legal text, cover letters and five-paragraph essays follow templates, and templates are predictable. So is writing by people with a smaller English vocabulary, as the Stanford study showed.
  • Mixed human and AI text. A draft where you wrote some paragraphs and AI wrote others is hard to score. Turnitin found that most of its sentence-level false positives sat right next to real AI writing, at the seams between the two.

If your own writing keeps getting flagged, why your writing is flagged as AI explains the usual causes.

Can AI detectors be wrong? How to use a score responsibly

They can, in both directions, as the evidence above shows. A score works best as one piece of evidence among several, never as the whole case.

If you're a student

  • Keep a record of how you wrote: drafts, notes, the sources you read, and version history in Google Docs or Word. That says more about authorship than any score.
  • Know your course's AI policy. If you used AI in a way it allows, say how you used it.
  • If your own work is flagged, ask which tool and score were used, point to the published false positive figures above, and offer to walk through your drafts or discuss the topic in person.
  • Checking a draft yourself shows which sentences read like AI to one detector. It can't tell you how another tool will score it, and it never replaces following your institution's rules.

If you're a teacher

  • Treat a score as a reason to look closer, never as the finding. Turnitin itself leaves that decision to you.
  • Compare with what you know: the student's earlier work, in-class writing, drafts, and whether they can talk through their argument.
  • Take extra care with short texts, low scores, non-native English writers and formulaic assignments, where false positives are most likely.
  • Set expectations before the assignment: which uses of AI are allowed, and what happens if a detector flags something.

Our page on AI checkers for teachers has more on using a checker in class.

AI detector accuracy FAQ

It depends on the text and the tool. In one 2023 test of 14 tools, none reached 80% accuracy, and results get worse on short, edited, translated or non-native English writing. Use a score as a signal, not proof.

AI checkers are AI detectors under another name, so the same limits apply. They do best on long, unedited AI drafts, worse on short or edited text, and they can flag human writing.

Yes, in both directions. They can flag human writing as AI (a false positive) and miss AI writing (a false negative). Even Turnitin says its false positive rate is not zero.

There's no honest single answer: results depend on the kind of text, and tools change with every update. We haven't published our own benchmark yet, so we don't rank detectors, ours included.

Detectors read predictable, simple word choices as a sign of AI, and writers with a smaller English vocabulary use more of them. A Stanford study found seven detectors labeled an average of 61% of TOEFL essays as AI.

Keep reading

Check a text, then read the score with care

Paste 40 to 200 words into the free AI detector to see which sentences read like AI. Treat the result as a signal, not proof.

Try the free AI detector