Methodology

How accurate is GPTOne? The numbers, the method, and the limits

GPTOne detects clean AI-generated text at 99.99% accuracy. This page is about the harder case. We ran 600 samples through five detectors under one threshold, deliberately loading the set with humanized rewrites, non-native English writing and mixed human and AI paragraphs, then published the raw counts. On that adversarial set GPTOne scored 94.4% overall with a 3.6% false positive rate, and missed 11 of 50 rewritten AI passages. That last number is the one we would want you to weigh most heavily.

This is GPTOne's own first-party study. It has not been independently replicated or peer-reviewed, and no detector on this market has been.

Key takeaways

  • 99.99% accuracy on clean AI-generated text. This is the figure quoted across the site. It comes from production evaluation on unmodified model output, not from the adversarial study below.
  • 94.4% overall accuracy on the 550-sample adversarial set, computed from the published raw counts. Lower by design: the set is built from the cases detectors fail on.
  • 3.6% false positive rate. 9 of 250 human-written samples were wrongly flagged as AI.
  • 8.0% on non-native English writing, 4 of 50. Double the general rate, and the metric we care about most.
  • 78% on humanized text, 39 of 50. Our weakest result, and the ceiling for the whole category.
  • Image accuracy has no single number. Images return a confidence score plus a region heatmap, never a fixed percentage.
600Samples in the study
99.99%Clean AI text, production
94.4%Overall, adversarial set
3.6%False positive rate
78%Humanized text caught
Method

What we tested, and how the dataset was built

Six hundred samples across ten subsets. Two hundred and fifty human, three hundred AI, fifty mixed. The mixed set is scored at sentence level and kept out of the binary accuracy figure, because averaging a document that is genuinely half-and-half produces a meaningless verdict.

SubsetSamplesLabelWhy it is in the study
Human, pre-AI era100HumanText that could not possibly be AI. The cleanest test of false positives there is.
Human, recent and verified100HumanModern human writing, which is harder than archival text because contemporary style has converged.
Human, non-native English50HumanThe subset where every detector on the market does worst, and the one with real consequences.
AI, GPT-550AIPer-model detection.
AI, Claude50AIPer-model detection.
AI, Gemini50AIPer-model detection.
AI, DeepSeek50AIPer-model detection.
AI, Llama50AIPer-model detection.
AI, humanized50AIAI output deliberately rewritten to evade detection. The hardest case in the set.
Mixed authorship50Sentence levelScored separately by sentence, not folded into the binary accuracy number.

Binary set for accuracy, false positive rate and false negative rate: 250 human plus 300 AI equals 550 samples. One AI-score threshold was chosen and applied identically to every tool, so the comparison is like for like even if your own preferred threshold differs.

Results

Five detectors, one threshold, the same 600 samples

Every figure below is derived from the raw per-subset counts we recorded, using the formulas stated above. Competitor figures come from running their publicly available tools under the same protocol at a single point in time.

ToolAccuracyFalse positivesFalse negativesFP, non-nativeHumanized caughtMixed, sentence level
GPTOne94.4%3.6%7.3%8.0%78%91.6%
Copyleaks88.9%8.0%13.7%22.0%60%82.0%
GPTZero82.9%12.4%21.0%34.0%44%78.0%
QuillBot80.0%12.8%26.0%30.0%40%74.0%
ZeroGPT77.8%16.8%26.7%42.0%36%71.0%

Lower is better for the two false-rate columns. Accuracy is ((250 − wrongly flagged humans) + correctly caught AI) / 550. False positive rate is wrongly flagged humans over 250. False negative rate is missed AI over 300.

Per-model detection rates

Each cell is out of 50 samples from that model, under the same threshold.

ToolGPT-5ClaudeGeminiDeepSeekLlama
GPTOne98%96%98%94%92%
Copyleaks96%92%94%90%86%
GPTZero94%88%86%82%80%
QuillBot88%82%84%78%72%
ZeroGPT90%84%80%76%74%

Every tool degrades in the same direction, from GPT-5 down to Llama. Detectors are tuned hardest on the model families that produced the most training text, and open-weight models are consistently the weakest link across the whole market.

Limitations

Six reasons not to read these numbers as certainty

1. We ran it, so we had an interest in the outcome

This is a first-party study. We designed the dataset, chose the threshold and recorded the counts. That is a genuine limitation and no amount of methodological care removes it. What we can do is state the method precisely enough that someone else could run it and get a different answer, and publish a result we lose on so the wins are worth something.

2. One threshold, applied everywhere

Every tool was scored at the same AI-score cutoff. That makes the comparison fair but it does not flatter any tool that ships a different default. Move the threshold and every number in both tables moves with it, typically trading false positives against false negatives.

3. Competitor products change, and this was a snapshot

These figures were measured at one point in time against publicly available versions of each tool. Detectors ship updates. A tool that did badly here may have improved since, and it is worth rechecking rather than citing an old table indefinitely.

4. Humanized text is the real ceiling

Our best result on deliberately rewritten AI output was 78%, which means 11 of 50 passages passed as human. Every tool in the study did worse. If your threat model is someone who paraphrases, detection is a weak control and process evidence is a strong one.

5. Short text produces unreliable scores

The samples in this study are full documents. Under roughly 200 words there is not enough text for the statistics to settle, and scores swing on small changes. Do not transfer these figures to a paragraph.

6. The false positive rate is not evenly distributed

The headline 3.6% conceals an 8.0% rate on non-native English writing. Second-language prose tends toward regular structure and conservative vocabulary, which is the same surface pattern a language model produces. Every tool in the study showed this effect, two of them at 34% and 42%. We wrote up the mechanism in why AI detectors falsely flag non-native English writers.

The line we will not cross

A detector score is a statistical measurement of text. It is not evidence of what a person did, it cannot be, and we will not describe it that way to sell the product. Anyone acting on a score, in a classroom or a hiring pipeline or an editorial workflow, needs a second source before they act.

Images

Why image accuracy has no single number

The figures above are text only. Image detection does not work the same way and should not be reported the same way.

A text sample arrives as text. An image arrives having been resized, re-encoded, cropped, screenshotted, or pushed through three messaging apps, and each of those steps destroys some of the signal a detector reads. An accuracy figure measured on clean generator output tells you almost nothing about a screenshot from a group chat. So GPTOne returns a confidence score for each image plus a region heatmap showing which areas of the picture drove it, rather than a fixed percentage.

Rather than assert a number, we published our method: a 60-image benchmark across five categories, mixing AI-generated pictures with real photographs so both catching fakes and not flagging real photos were measured. That is our own first-party testing, not a third-party evaluation. One concrete result from it: a photoreal AI-generated portrait, the kind that passes a quick human glance, was flagged as AI Generated at 97% confidence with a High confidence badge.

The write-up, and what to make of any image detector's accuracy claim, is in how accurate are AI image detectors, tested on 60 images and what an AI image confidence score actually means.

Common questions

Questions about how we measure accuracy

No. It is GPTOne's own first-party benchmark, designed and run by us, and you should read it with that in mind. What we can offer instead of independence is transparency: the dataset composition, the formulas, the per-tool raw counts and the single threshold we applied to every tool are all stated, so the figures can be recomputed and the method can be repeated by someone with no stake in the result. No detector on this market, ours included, has been independently peer-reviewed.
Because a study where one tool wins everything is not a study, it is an advertisement. Our weakest result is humanized text, where we caught 39 of 50, a 78% detection rate. That is the best figure in the comparison and it still means roughly one rewritten passage in five gets through. If you are relying on detection to catch students or contractors who paraphrase, that is the number you should be planning around.
There is no single percentage, and any tool quoting one for images is oversimplifying. GPTOne returns a confidence score for the image plus a region heatmap showing which areas drove it. Compression, cropping, screenshotting and re-encoding all move the result, so a figure measured on clean files does not transfer to a photo that has been through three messaging apps. We published our method and results from a 60-image, 5-category first-party test rather than asserting a number.
It means 9 of the 250 human-written samples in our study were flagged as AI. Scaled to a 120-paper cohort, that is roughly four papers that would be wrong. The rate is also not evenly distributed: on the 50 samples by non-native English writers it was 8.0%. Any process that treats a flag as a finding will, at that rate, produce wrong findings, and it will produce them disproportionately against second-language writers.
Partly. Detection accuracy is a moving target: every new model release shifts the writing patterns a classifier was tuned on, and every humanizer release attacks the signal directly. The metrics most likely to drift are per-model detection and humanized detection. The false positive rate is the more stable number, and it is the one we would judge a detector on first.

The reliability questions behind these numbers

Longer treatments of the three things people get wrong about detector accuracy.

Check the numbers against your own text

A published benchmark is worth something. Your own documents are worth more. Sign up free and your starting credits cover your first scans.

Run a free AI scan

Free credits on signup · No card required · Up to 50,000 characters per scan on every plan