Paste a passage and GPTOne returns a 0 to 100 score, a human, mixed or AI band, and the sentences that pushed the score. On our own 600-sample benchmark it caught 49 of 50 GPT-5 samples. It also gets things wrong sometimes, and this page is about both halves of that.
No detector reads a hidden marker. ChatGPT does not stamp its output. What a detector does is compare your text against the statistical shape of text a language model tends to produce, and report how close the match is.
Two properties carry most of the signal.
A language model picks words that are, on average, the likeliest continuation. That makes its output low-perplexity: each word is close to what a model would have predicted. Human writing wanders. We pick the odd word, double back, use a phrase that does not quite fit. Sustained low perplexity across several hundred words is the strongest single tell.
People write a long sentence, then a short one. Then a fragment. Model output is more even, with sentence lengths and structures clustered tightly around a mean. Low burstiness on its own proves nothing, plenty of humans write evenly, but low burstiness plus low perplexity is a much stronger combination than either alone. There is a longer walkthrough in perplexity and burstiness explained.
GPTOne scores every sentence and returns the highest-scoring ones alongside the document score. This matters more than the headline figure. A document at 64 where three consecutive paragraphs sit above 90 and the rest sit below 20 is a different situation from a document at 64 that is uniformly grey, and only the sentence view tells you which one you are holding.
These are the same thresholds the detection API returns, so what you see on the web and what a developer gets back from an integration agree.
| Score | Band | What it usually means | Sensible next step |
|---|---|---|---|
| 0 to 50 | Human | The text behaves like human writing across both perplexity and rhythm. | Nothing further, unless you have a separate reason to look. |
| 51 to 60 | Mixed | Often genuinely mixed authorship, or human text that has been heavily edited for smoothness. | Read the sentence breakdown. A cluster tells a different story from an even spread. |
| 61 to 100 | AI | The text sits well inside the machine-written range. | Still a signal, not a finding. Look at which sentences drove it and ask the writer. |
Every band is probabilistic. A high score on a 60-word paragraph is far weaker evidence than the same score on 2,000 words, because short texts do not give the classifier enough to work with.
The most-detected output on the internet. Both models are strongly patterned and most classifiers, ours included, were trained on plenty of their text.
StrongMore conversational and better at varying rhythm, so burstiness alone helps less. The perplexity signal still holds up.
StrongThe reasoning models write longer and more structured answers. Structure is itself a signal, though it overlaps with careful human academic writing, which is where false positives come from.
Good49 of 50 samples flagged in our benchmark. Newer models are generally harder because they write with more variation, and this is the number worth rechecking as releases land.
98% in our studyIf you need Claude and Gemini in the same pass rather than ChatGPT alone, the multi-model detector page covers all three families. And do AI detectors work on GPT-5 goes deeper on why each new release resets part of the problem.
These are the four conditions that degrade a result most, in rough order of how often they show up.
This is the big one. On the 50 humanized samples in our benchmark GPTOne caught 39, a 78% rate. That is the best figure in the study and it is still one miss in five. Paraphrasing raises perplexity and breaks the rhythm signature, which is exactly what the classifier reads.
Under roughly 200 words there is not enough text for the statistics to settle. Short-text scores swing hard on small changes, so a high score on a paragraph should carry far less weight than the same score on a full essay.
Second-language writing tends toward simpler, more regular constructions, which is the same surface pattern a model produces. GPTOne wrongly flagged 4 of 50 non-native samples, an 8% rate. Better than the 34% and 42% we measured for two other tools on the same set, but still double our overall false positive rate. We wrote about the mechanism in why AI detectors falsely flag non-native English writers.
A single document score averages across passages that deserve different verdicts. This is where sentence-level results stop being a nice extra and become the only useful output.
A detector score is evidence of a statistical pattern. It is not evidence of what a person did. Anyone acting on a score, in a classroom or a newsroom or a hiring pipeline, needs a second source before they act.
Three pieces that cover the parts this page only summarises.
Sign up free and your starting credits cover your first scans. Up to 50,000 characters at a time, with sentence-level results rather than one number.
Run a free AI scanFree credits on signup · No card required · ChatGPT, Claude, Gemini, Grok, DeepSeek and LLaMA