GPTOne detects clean AI-generated text at 99.99% accuracy. This page is about the harder case. We ran 600 samples through five detectors under one threshold, deliberately loading the set with humanized rewrites, non-native English writing and mixed human and AI paragraphs, then published the raw counts. On that adversarial set GPTOne scored 94.4% overall with a 3.6% false positive rate, and missed 11 of 50 rewritten AI passages. That last number is the one we would want you to weigh most heavily.
Six hundred samples across ten subsets. Two hundred and fifty human, three hundred AI, fifty mixed. The mixed set is scored at sentence level and kept out of the binary accuracy figure, because averaging a document that is genuinely half-and-half produces a meaningless verdict.
| Subset | Samples | Label | Why it is in the study |
|---|---|---|---|
| Human, pre-AI era | 100 | Human | Text that could not possibly be AI. The cleanest test of false positives there is. |
| Human, recent and verified | 100 | Human | Modern human writing, which is harder than archival text because contemporary style has converged. |
| Human, non-native English | 50 | Human | The subset where every detector on the market does worst, and the one with real consequences. |
| AI, GPT-5 | 50 | AI | Per-model detection. |
| AI, Claude | 50 | AI | Per-model detection. |
| AI, Gemini | 50 | AI | Per-model detection. |
| AI, DeepSeek | 50 | AI | Per-model detection. |
| AI, Llama | 50 | AI | Per-model detection. |
| AI, humanized | 50 | AI | AI output deliberately rewritten to evade detection. The hardest case in the set. |
| Mixed authorship | 50 | Sentence level | Scored separately by sentence, not folded into the binary accuracy number. |
Binary set for accuracy, false positive rate and false negative rate: 250 human plus 300 AI equals 550 samples. One AI-score threshold was chosen and applied identically to every tool, so the comparison is like for like even if your own preferred threshold differs.
Every figure below is derived from the raw per-subset counts we recorded, using the formulas stated above. Competitor figures come from running their publicly available tools under the same protocol at a single point in time.
| Tool | Accuracy | False positives | False negatives | FP, non-native | Humanized caught | Mixed, sentence level |
|---|---|---|---|---|---|---|
| GPTOne | 94.4% | 3.6% | 7.3% | 8.0% | 78% | 91.6% |
| Copyleaks | 88.9% | 8.0% | 13.7% | 22.0% | 60% | 82.0% |
| GPTZero | 82.9% | 12.4% | 21.0% | 34.0% | 44% | 78.0% |
| QuillBot | 80.0% | 12.8% | 26.0% | 30.0% | 40% | 74.0% |
| ZeroGPT | 77.8% | 16.8% | 26.7% | 42.0% | 36% | 71.0% |
Lower is better for the two false-rate columns. Accuracy is ((250 − wrongly flagged humans) + correctly caught AI) / 550. False positive rate is wrongly flagged humans over 250. False negative rate is missed AI over 300.
Each cell is out of 50 samples from that model, under the same threshold.
| Tool | GPT-5 | Claude | Gemini | DeepSeek | Llama |
|---|---|---|---|---|---|
| GPTOne | 98% | 96% | 98% | 94% | 92% |
| Copyleaks | 96% | 92% | 94% | 90% | 86% |
| GPTZero | 94% | 88% | 86% | 82% | 80% |
| QuillBot | 88% | 82% | 84% | 78% | 72% |
| ZeroGPT | 90% | 84% | 80% | 76% | 74% |
Every tool degrades in the same direction, from GPT-5 down to Llama. Detectors are tuned hardest on the model families that produced the most training text, and open-weight models are consistently the weakest link across the whole market.
This is a first-party study. We designed the dataset, chose the threshold and recorded the counts. That is a genuine limitation and no amount of methodological care removes it. What we can do is state the method precisely enough that someone else could run it and get a different answer, and publish a result we lose on so the wins are worth something.
Every tool was scored at the same AI-score cutoff. That makes the comparison fair but it does not flatter any tool that ships a different default. Move the threshold and every number in both tables moves with it, typically trading false positives against false negatives.
These figures were measured at one point in time against publicly available versions of each tool. Detectors ship updates. A tool that did badly here may have improved since, and it is worth rechecking rather than citing an old table indefinitely.
Our best result on deliberately rewritten AI output was 78%, which means 11 of 50 passages passed as human. Every tool in the study did worse. If your threat model is someone who paraphrases, detection is a weak control and process evidence is a strong one.
The samples in this study are full documents. Under roughly 200 words there is not enough text for the statistics to settle, and scores swing on small changes. Do not transfer these figures to a paragraph.
The headline 3.6% conceals an 8.0% rate on non-native English writing. Second-language prose tends toward regular structure and conservative vocabulary, which is the same surface pattern a language model produces. Every tool in the study showed this effect, two of them at 34% and 42%. We wrote up the mechanism in why AI detectors falsely flag non-native English writers.
A detector score is a statistical measurement of text. It is not evidence of what a person did, it cannot be, and we will not describe it that way to sell the product. Anyone acting on a score, in a classroom or a hiring pipeline or an editorial workflow, needs a second source before they act.
The figures above are text only. Image detection does not work the same way and should not be reported the same way.
A text sample arrives as text. An image arrives having been resized, re-encoded, cropped, screenshotted, or pushed through three messaging apps, and each of those steps destroys some of the signal a detector reads. An accuracy figure measured on clean generator output tells you almost nothing about a screenshot from a group chat. So GPTOne returns a confidence score for each image plus a region heatmap showing which areas of the picture drove it, rather than a fixed percentage.
Rather than assert a number, we published our method: a 60-image benchmark across five categories, mixing AI-generated pictures with real photographs so both catching fakes and not flagging real photos were measured. That is our own first-party testing, not a third-party evaluation. One concrete result from it: a photoreal AI-generated portrait, the kind that passes a quick human glance, was flagged as AI Generated at 97% confidence with a High confidence badge.
The write-up, and what to make of any image detector's accuracy claim, is in how accurate are AI image detectors, tested on 60 images and what an AI image confidence score actually means.
Longer treatments of the three things people get wrong about detector accuracy.
A published benchmark is worth something. Your own documents are worth more. Sign up free and your starting credits cover your first scans.
Run a free AI scanFree credits on signup · No card required · Up to 50,000 characters per scan on every plan