AI Detection Reliability Study: The False Positives Problem in 2025
Sana Bano
·May 18, 2026
·8 min read
Three peer-reviewed studies on AI detector false positives. Liang et al. found over 50% of non-native English essays wrongly flagged as AI. The evidence.
Three peer-reviewed studies have tested whether AI text detectors produce false positives, and all three found that they do. The largest effect is on non-native English writers: Liang et al. found detectors misclassified more than half of non-native TOEFL essays as AI-generated while scoring near-perfect on US 8th-grade essays. Below is what the published evidence actually says, and what it means if you have to act on a detector score.
Key Takeaways
- Liang et al. (2023, Patterns) found that over 50% of TOEFL essays by non-native English writers were flagged as AI-generated, against near-perfect accuracy on US 8th-grade essays written by native speakers.
- Weber-Wulff et al. (2023, International Journal for Educational Integrity) tested 16 detection tools and found that 10 of the 16 classified ChatGPT-generated text as human-written, and that paraphrasing significantly reduced accuracy for five of them.
- Sadasivan et al. (2023) proved a theoretical impossibility result: as language models get better at imitating human writing, the best possible detector approaches the performance of a random classifier.
- In our own 600-sample benchmark, GPTOne produced a 3.6% false positive rate overall and 8.0% on non-native English writing. That non-native rate is more than double the general rate, which matches the direction of the published findings.
- No detector on the market has published an independently replicated false-positive rate. Every accuracy number in this space, including ours, is vendor-run or academic-run on a fixed sample.
What counts as a false positive
A false positive is human-written text that a detector labels as AI-generated. It is the error that carries consequences. A missed AI passage costs a grade or a byline. A false positive puts a real person in front of an academic integrity board over work they actually wrote.
The asymmetry matters for how you read any accuracy claim. A detector can post a high overall accuracy number while still failing badly on one group of writers, because that group is a small share of the test set. Overall accuracy hides distribution. False positive rate by writer population is the number that tells you whether a tool is safe to act on.
The published evidence
Liang et al., 2023: the non-native English bias
The clearest finding in the literature comes from Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou at Stanford, published in Patterns. They ran seven widely used GPT detectors over TOEFL essays written by non-native English speakers and over essays written by US 8th-grade students.
The detectors classified more than half of the non-native essays as AI-generated. On the US 8th-grade essays, they were close to perfect.
The authors' explanation is mechanical rather than malicious. Detectors lean on perplexity, a measure of how predictable the next word is. Writers with a smaller working vocabulary and less syntactic variety produce text with lower perplexity, which is the same signal a language model produces. The researchers tested this directly: enriching the word choice in non-native samples reduced misclassification, and simplifying the native samples increased it. The detector is not identifying AI. It is identifying linguistic simplicity and treating it as a proxy.
Read the paper: GPT detectors are biased against non-native English writers00130-7), Patterns, 2023. Preprint at arXiv:2304.02819.
Weber-Wulff et al., 2023: most tools miss AI text entirely
Debora Weber-Wulff and colleagues tested 16 detection tools for the International Journal for Educational Integrity. Their finding runs in the opposite direction from the bias result and is just as damaging.
Human-written text was correctly identified as human by all the tools they tested. But ChatGPT-generated text was predicted to be human-written by 10 of the 16. Applying paraphrasing to AI text significantly lowered detection accuracy for five of the tools.
Put the two studies side by side and you get the real picture. Detectors are not uniformly trigger-happy or uniformly lax. They are unreliable in both directions at once, and which direction depends on who wrote the text and whether it was edited after generation.
Read the paper: Testing of detection tools for AI-generated text, International Journal for Educational Integrity, 2023.
Sadasivan et al., 2023: the ceiling is mathematical
Vinu Sankar Sadasivan and co-authors went further and asked whether reliable detection is possible even in principle. Their answer, both empirically and theoretically, is that it is not, at the limit.
Empirically, they showed that a lightweight paraphraser applied on top of a language model breaks a wide range of detectors, including watermarking schemes, neural classifiers and zero-shot methods. Retrieval-based detectors built specifically to resist paraphrasing still fell to recursive paraphrasing.
Theoretically, they connect the best achievable detector performance to the statistical distance between human and AI text distributions. As models improve at imitating human writing, that distance shrinks, and the best possible detector converges toward a coin flip.
This is the part most vendor marketing omits. There is a ceiling, it is not an engineering problem, and no amount of training data moves it.
Read the paper: Can AI-Generated Text be Reliably Detected?, 2023.
Our own measurement, and how it compares
We ran a 600-sample benchmark across five detectors under a single threshold and published the raw counts so the figures can be checked. This is GPTOne's own first-party study. It has not been independently replicated or peer-reviewed, and neither has any competitor's.
On the 550-sample binary set, measured on identical inputs:
| Detector | False positive rate | FPR on non-native English |
|---|---|---|
| GPTOne | 3.6% | 8.0% |
| Copyleaks | 8.0% | 22.0% |
| GPTZero | 12.4% | 34.0% |
| QuillBot | 12.8% | 30.0% |
| ZeroGPT | 16.8% | 42.0% |
Two things are worth saying plainly about our own column.
First, the direction of the Liang finding reproduces in our data. Our non-native false positive rate, 8.0%, is more than double our general rate of 3.6%. We have the lowest figure in the group and it is still the number we are least comfortable with.
Second, 8.0% means roughly one non-native writer in twelve gets wrongly flagged. That is not a rate at which anyone should be making an accusation from a single score.
We go deeper on the mechanism in why AI detectors falsely flag non-native English, and on tool selection in which AI detector has the lowest false positive rate.
What this means if you have to act on a score
The evidence supports a narrow set of conclusions and does not support the broad ones.
A score is a signal for review, never proof. No published study supports treating a detector output as evidence of misconduct on its own. The Liang result alone makes a single-score accusation indefensible against any non-native English writer.
Scan the whole document, not a fragment. Detectors read patterns across a passage. Short samples produce unstable scores in both directions.
Weight non-native writing differently, or exclude it from automated flagging. If your cohort includes international students or ESL writers, an automated threshold applied uniformly will concentrate its errors on them.
Treat an edited or paraphrased document as undetectable. Sadasivan et al. and Weber-Wulff et al. agree on this from different angles. If the text has been through a rewriting tool, a clean score tells you nothing.
Ask any vendor for false positive rate by writer population. Overall accuracy is close to meaningless here. If a vendor will not publish the non-native figure, assume it is bad.
If you need to check your own writing before submitting it, our AI content detector reports which passages carry AI signal rather than a bare percentage, which is the part you can actually argue with. For students facing a flag, how to prove you did not use AI on an essay covers the documentation that holds up.
The limits of this evidence
Being straight about what this body of work does not settle:
- The studies are from 2023. Detectors and language models have both moved since. The direction of the bias finding is unlikely to have reversed, because the mechanism is perplexity itself, but the magnitudes may differ today.
- Sample sizes are modest. Liang et al. used 91 TOEFL essays and 88 US 8th-grade essays. That is enough to establish a large effect, not enough to pin down a precise rate.
- No study covers every current model. Coverage of Claude, Gemini and the newer GPT releases is thin in the peer-reviewed literature.
- Our own numbers are first-party. We ran them, we published the raw counts, and that is not the same as independent replication.
- Nobody has published a longitudinal false positive rate measured on real submissions rather than a constructed test set. That is the study the field actually needs.
FAQ
Is there a peer-reviewed study on AI detector false positives?
Yes. The most cited is Liang et al. (2023) in Patterns, which found detectors misclassified over half of non-native English TOEFL essays as AI-generated. Weber-Wulff et al. (2023) in the International Journal for Educational Integrity tested 16 tools and found most failed in the opposite direction.
What is a normal false positive rate for an AI detector?
Published and vendor-run figures range from roughly 3% to 17% on general text, and from 8% to 42% on non-native English writing. Anyone quoting a single rate without saying which population it was measured on is not telling you much.
Can a detector prove a student used AI?
No. No peer-reviewed study supports that use. Detector output is a signal that justifies a conversation, not evidence that supports a finding.
Why do detectors flag non-native English writers more often?
Because they use perplexity as a proxy for machine generation. Writing with less lexical variety scores as more predictable, and predictable is what the model looks for. Liang et al. confirmed this by enriching vocabulary in non-native samples and watching the misclassifications fall.
Are detectors getting more reliable over time?
Not necessarily. Sadasivan et al. showed that as language models improve at imitating human text, the theoretical ceiling on detector performance falls. Better models make the problem harder, not easier.
Detector scores are useful for deciding where to look. They are not useful for deciding what happened. If you want to see the reasoning rather than a number, run the text through our free AI detector and read the flagged passages, not the score.