Does GPTZero Detect Claude? What Its Own Test Actually Shows
Sana Bano
·September 13, 2026
·9 min read
Yes, and GPTZero published a test claiming a 100% catch rate on Claude across 30 samples. Here is what that evidence is actually worth.
Yes. GPTZero detects Claude text and publishes a study saying it caught all 30 of its test samples across Claude Opus 4.8, Sonnet 4.6 and Haiku 4.5. That is a real result. It is also a vendor grading its own homework on 30 samples, and the author says so openly.
The number is less interesting than what sits behind it.
Key Takeaways
- GPTZero tested three Claude models: Opus 4.8, Sonnet 4.6 and Haiku 4.5, reporting a 100% catch rate across 30 total samples.
- The sample size is 30. GPTZero's own author calls the testing basic and the sample small.
- Controlled benchmarks run high. GPTZero reports 99.5% on a University of Chicago Booth benchmark from February 2026 covering GPT-4.1, Claude Opus 4, Claude Sonnet 4 and Gemini 2.0 Flash.
- Real-world performance runs lower. Third-party reviewers consistently report figures in the low-to-mid 80s, though none publish their test sets.
- Hybrid text is the hard case. Human-edited machine output defeats detectors far more reliably than raw generation does.
What GPTZero actually published
GPTZero ran the test itself and put it on its own news page. You can read GPTZero's Claude test directly.
The setup: five content types, three Claude models, two phases, 30 samples total. Reported per-model detection ran from 87.2% to 100% depending on model and phase, with every sample ultimately receiving an AI verdict. The write-up concludes that Claude fooled other detectors but could not fool GPTZero once.
Credit where it is due on two points. GPTZero tested against current Claude models rather than a two-year-old snapshot, which many competitors do not bother with. And the author states plainly that the sample size was small and the testing basic.
That honesty is worth more than the percentage.
Why 30 samples cannot settle this
Here is the statistical problem, and it applies to every vendor study of this shape.
With 30 samples, the confidence interval around a 100% result is wide. A tool that genuinely catches 90% of Claude output will return a perfect score on 30 samples reasonably often, purely by chance. You cannot distinguish 100% from 90% at that sample size, and the difference between those two numbers is enormous when applied across a university's annual submissions.
The second problem is selection. The samples were chosen by the party being tested. Not dishonestly, but inevitably: you generate text the way you expect people to generate it, and that expectation shapes what you produce. Real student writing arrives edited, mixed, truncated and pasted together.
The third problem is the one that actually matters, and no vendor study addresses it.
The number nobody publishes
Detection rate answers "how often does this catch AI text". It says nothing about how often the tool accuses a human.
A detector tuned aggressively enough will catch 100% of Claude output. It will also flag a meaningful share of human essays, and that cost falls unevenly. Stanford-led research in Liang et al., 2023 found detectors misclassifying writing by non-native English speakers as AI-generated at dramatically higher rates than writing by native speakers, because restricted vocabulary and even sentence rhythm read as machine-like to a statistical classifier.
A 100% catch rate published without its matching false-positive rate, measured on the same set, is half an answer. It is the flattering half.
This is why we report both figures together from our own 600-sample benchmark, and why our comparison of which AI detector has the lowest false positive rate leads with the error side rather than the accuracy side.
Why Claude is genuinely harder than ChatGPT
GPTZero's result is more impressive than it first appears, because Claude is a harder target than earlier model generations.
Detectors read two properties. Perplexity measures how predictable each word is given what came before. Burstiness measures how much sentence length and complexity vary across a passage. Machine text has historically been predictable and even. Human writing is less predictable and lumpier. We explained the mechanics in how AI detectors work.
Claude scores closer to human on both. Its vocabulary range is wider, its sentence rhythm more varied, and its default register more qualified and hedged. Those are the characteristics of considered human prose, which is precisely why writers prefer it and precisely why it sits nearer the decision boundary.
So a detector that handles Claude well is doing real work. A detector last calibrated against GPT-3.5 output is measuring a statistical profile current models no longer produce, which is the argument we made in do AI detectors need Claude and Gemini coverage to be reliable.
The benchmark gap
GPTZero reports 99.5% on a controlled academic benchmark. Third-party reviewers writing about real-world use consistently land lower, often in the low-to-mid 80s.
Both can be true, and the gap is not evidence of dishonesty.
Benchmark text is clean: full-length, monolingual, generated in one pass, unedited. Real text is messy: short, edited, translated, pasted from three sources, written by someone working in a second language. Every one of those conditions degrades detector performance, and none of them appear in a curated evaluation set.
The exception worth knowing about is the RAID benchmark, built specifically to test detectors against adversarial pressure like paraphrasing and synonym substitution. Results there scatter widely, and rankings change depending on which transformation is applied. That is the honest picture of this category.
A practical consequence: treat any vendor accuracy figure as an upper bound on a good day, not as the number you will see on your document.
What to do if you need a real answer
Three steps.
Check the length. Below roughly 300 words, no detector on the market gives a dependable answer, GPTZero included. Short passages do not carry enough signal.
Run a second, independent tool. Agreement between two detectors is far more informative than one high number. Disagreement tells you the text sits in the ambiguous band where no percentage should drive a decision. Our free AI detector covers Claude, ChatGPT, Gemini, GPT-5, Grok, DeepSeek and LLaMA at 99.99% accuracy on text, with free credits on signup, no card required, and up to 50,000 characters per scan on every plan including Free. For a like-for-like look at the two tools, see GPTOne versus GPTZero.
Never let a score be the only evidence. Not ours, not GPTZero's. Detection is probabilistic, the output is a likelihood, and the consequences of treating it as a finding of fact land on real people.
If you want the deeper treatment of detecting Claude specifically, we wrote a full guide to Claude AI detection, and the same question about a rival tool in does ZeroGPT detect Claude.
What a study worth trusting would look like
None of this means vendor testing is worthless. It means the bar is low across the whole category, and it is worth knowing what a better study looks like so you can spot one.
Four things would move GPTZero's Claude test from encouraging to conclusive.
A bigger sample. Several hundred documents per model rather than ten, so the confidence interval narrows enough to distinguish 100% from 92%.
A human control group. Run the same number of genuine human documents through the same pipeline and report how many got flagged. Without this, a detection rate is uninterpretable.
Messy inputs. Edited text, translated text, short passages, mixed human and machine paragraphs. The clean generated sample is the easiest case and the least representative one.
A published test set. So somebody else can check. This is the single thing that separates a benchmark from a marketing claim, and it is why a result on a shared academic set carries weight that a private one does not.
Until a vendor does all four, treat every accuracy figure in this market, ours included, as a claim rather than a measurement.
FAQ
Does GPTZero detect Claude?
Yes. GPTZero published a test across Claude Opus 4.8, Sonnet 4.6 and Haiku 4.5 reporting a 100% catch rate on 30 samples.
Is a 100% catch rate reliable?
On 30 self-selected samples, no. The confidence interval is wide and the study does not report a false-positive rate on human writing.
Why is Claude harder to detect than ChatGPT?
Claude produces wider lexical variety and more varied sentence rhythm, both of which push statistical detectors toward a human classification.
Can GPTZero catch Claude text that a human edited?
Hybrid writing is the hardest case for every detector in this category. Editing disrupts the statistical patterns classifiers rely on.
What does GPTZero cost?
Its published pricing starts at Premium, with a free entry point whose limits are no longer listed in the comparison table. We broke this down in GPTZero's free tier limits.
The short version
GPTZero does detect Claude, and it did the work of testing against current models. The evidence is a self-run 30-sample study with no false-positive figure attached, so read it as encouraging rather than conclusive.
Get a second opinion at GPTOne. Free credits on signup, no card required.
Meta description: Yes, and GPTZero published a test claiming a 100% catch rate on Claude across 30 samples. Here is what that evidence is actually worth.