Is GPTZero accurate? The short answer
GPTZero is accurate on plain, unedited AI text in the most recent independent tests, and it rarely flagged human writing in them. It was much less reliable in 2023, when researchers found it labeling essays by non-native English writers as AI. It still struggles with AI text that has been reworked to sound human or written in a real author's style.
We have not tested GPTZero ourselves. Everything below comes from GPTZero's own published claims and from independent studies, each one linked. If you want to check a text right now, our free AI detector shows which sentences read like AI, but it's a different model, so its score won't match GPTZero's.
What is GPTZero?
GPTZero (sometimes written GPT Zero) is an AI detector made by Edward Tian, who released it on January 2, 2023 while he was a senior at Princeton. Back then it judged text by "perplexity" and "burstiness": how predictable the words are, and how much the sentences vary. Its team page lists Tian as co-founder and CEO and Alex Cui as co-founder and CTO.
Today GPTZero describes its detector as an end-to-end deep learning model with a sentence-by-sentence classifier. It labels a document as written entirely by a human, entirely by AI, or a mix of the two.
How accurate is GPTZero, by its own numbers?
GPTZero publishes its own accuracy figures. They come from the company's benchmarks, on test sets it chose, so read them as claims.
- Technology page (undated, read October 3, 2026). GPTZero says it keeps its false positive rate "at no more than 1%" when evaluating AI versus human text. It reports 96.5% accuracy on documents that mix human and AI writing, and says it cut its false positive rate on TOEFL essays to 1.1%.
- GPTZero 4o (announced September 24, 2026). In the launch post for the model it made the default on September 20, 2026, GPTZero reports 1 false positive out of 10,402 student essays. It also claims over 99% recall on GPT-5.6, Gemini 3.6, Grok 4.5 and Claude 5, at a false positive rate under 0.03%, on its own benchmark.
- Its own caveat. The technology page also says "No AI detector is 100% accurate", that GPTZero does best on longer texts and English prose, and that it should be used "as a conversation starter, and not as the final verdict".
A company's numbers on its own test set tell you how the tool performs on that kind of text. They can't tell you how it will handle yours.
What independent studies found
Researchers outside the company have tested GPTZero several times. The results changed a lot between 2023 and 2026, and so did GPTZero's model, so check the date of each test before you quote it.
| Study | When GPTZero was tested | What it found for GPTZero |
|---|---|---|
| Liang et al., Stanford, published in Patterns | March 2023 | Flagged 52% of 91 TOEFL essays by non-native English writers as AI, and none of 88 essays by US eighth graders |
| Weber-Wulff et al., 14 tools | March 2023 | Flagged 9 of 18 human texts as AI, 6 of them machine-translated into English; missed fewer AI texts than any other tool |
| Pratama, PeerJ Computer Science | Late December 2024 | 97.22% accuracy on scholarly abstracts; flagged none of 72 human-written abstracts, half of them by non-native English writers |
| Jabarian and Imas, Chicago Booth working paper | 2025 (paper dated September 2025) | Flagged about 0.7% of human passages; missed 0.2% to 3% of plain AI passages; missed about half or more of AI text run through a humanizer |
| Epoch AI | June 2026 | Flagged 0 of 495 human passages; missed 2 of 297 plain AI passages and 32 of 297 written in a real author's style |
The 2023 studies
Liang's team traced the false positives to vocabulary. When they had ChatGPT enrich the word choices in the TOEFL essays, GPTZero's rate fell from 52% to 19%. When they simplified the vocabulary of the eighth-grade essays, it rose from 0% to 39%. GPTZero caught all 31 ChatGPT-written college essays they gave it, but only 13% once they asked ChatGPT to rewrite them in more literary language.
Weber-Wulff's team was blunt. Every other tool in their test did better on human-written text than on AI text; GPTZero was the exception. They wrote that "half of the positive classifications would be false accusations, which makes this tool unsuitable for the academic environment."
The 2025 and 2026 studies
The newer tests look very different. Pratama's study of scholarly abstracts found no false positives for GPTZero on abstracts by native or non-native English writers.
The Chicago Booth paper tested human passages from six genres, from news and blogs to novels and résumés, against text from GPT-4.1, Claude Opus 4, Claude Sonnet 4 and Gemini 2.0 Flash. The authors ranked Pangram first and put GPTZero in a second tier with Originality.ai. The paper is a working paper, not yet peer-reviewed.
Epoch AI ran GPTZero's 2026-05-11-base model in June 2026 on passages of about 500 words. It flagged none of 495 human passages from blogs, fiction and science written before 2022. Epoch counted a "mixed" verdict as a miss, which makes its miss rates stricter than a simple yes or no.
Where GPTZero struggles
- Reworded or humanized AI text. In the Chicago Booth paper, GPTZero missed about half or more of AI passages run through the humanizer StealthGPT. In Liang's study, a literary-language rewrite cut its catch rate from 100% to 13%.
- AI writing in a real author's style. Epoch AI found GPTZero missed 10.8% of passages where the AI imitated a real author, and 24% in scientific writing, against 0.67% for basic prompts.
- Short text. The Booth paper found GPTZero weaker on snippets under 50 words, and GPTZero itself says it does best on longer texts.
- Non-native English, in older versions. The 2023 results on TOEFL essays were poor. Newer evidence, including Pratama's study and GPTZero's own 1.1% claim, points the other way, but one study of 72 abstracts doesn't settle it.
- Famous, much-copied text. In 2023, Ars Technica covered a viral screenshot in which GPTZero called a section of the US Constitution "likely to be written entirely by AI". Tian's explanation was that the Constitution shows up again and again in the data these models learn from.
A missed AI text and a false accusation aren't equal problems. Teachers worry about the first. Students pay for the second.
GPTZero vs ZeroGPT: different products, different companies
The names are near mirror images, and they're easy to mix up. They are separate detectors. GPTZero lives at gptzero.me and was started by Edward Tian. ZeroGPT lives at zerogpt.com, and its terms of use name Olive Works LLC, a Wyoming company, as the business behind it.
ZeroGPT is the one we have tested. On October 2, 2026, we ran 40 texts through it. It caught all 20 AI texts and cleared all 14 modern human texts, but scored six famous older texts, the US Constitution and the Gettysburg Address among them, 73 to 100% AI. The method and every score are in is ZeroGPT accurate. None of those numbers apply to GPTZero.
Few studies run both tools on the same texts. Pratama's is one: on the same scholarly abstracts, GPTZero reached 97.22% accuracy and ZeroGPT 64.35%.
How to read a GPTZero score fairly
- Read the label, not just the number. GPTZero sorts text into human, AI or mixed and gives a confidence level. It reports an average error rate under 1% for its "high" confidence predictions, so treat lower-confidence results with more caution.
- Treat "mixed" as exactly that. A document can be partly yours and partly AI-assisted. A mixed result says so and nothing more.
- Weigh length and editing. Short texts and heavily edited ones are where the studies above found GPTZero weakest.
- Use it as one piece of evidence. Compare it with drafts, version history and in-class writing. GPTZero itself calls its result a conversation starter.
- Expect detectors to disagree. Tools use different models and thresholds. Our guide on whether AI detectors are accurate explains why, and what false positives mean in practice.
- If your own writing is flagged, read why your writing gets flagged as AI for the usual causes and the evidence that helps.
If you use AI to draft where your course allows it, our GPTZero humanizer page explains what a rewrite changes and what no tool can promise. We haven't tested any humanizer against GPTZero, and your school's rules come first.