ChatGPT detection

ChatGPT detector: check whether text came out of ChatGPT

Paste a passage and GPTOne returns a 0 to 100 score, a human, mixed or AI band, and the sentences that pushed the score. On our own 600-sample benchmark it caught 49 of 50 GPT-5 samples. It also gets things wrong sometimes, and this page is about both halves of that.

Free credits on signup, no card required. Up to 50,000 characters per scan, about 8,000 words, on every plan.

Key takeaways

  • 98% detection on clean GPT-5 output. 49 of 50 samples flagged correctly in GPTOne's own 600-sample study.
  • 78% on humanized text. Rewrite ChatGPT output by hand and roughly one passage in five slips through. That is the honest ceiling of the whole category.
  • 3.6% false positive rate. 9 of 250 human-written samples were wrongly flagged as AI. Low, but not zero.
  • Score bands are 0 to 50 human, 51 to 60 mixed, 61 to 100 AI. Same thresholds the API returns.
  • 50,000 characters per scan, about 8,000 words, on Free and paid alike. Scanning requires a free account.
98%GPT-5 samples caught (49/50)
3.6%False positive rate (9/250)
78%Humanized text caught (39/50)
50,000Characters per scan, all plans
The mechanism

What a ChatGPT detector is actually measuring

No detector reads a hidden marker. ChatGPT does not stamp its output. What a detector does is compare your text against the statistical shape of text a language model tends to produce, and report how close the match is.

Two properties carry most of the signal.

Perplexity: how surprising the next word is

A language model picks words that are, on average, the likeliest continuation. That makes its output low-perplexity: each word is close to what a model would have predicted. Human writing wanders. We pick the odd word, double back, use a phrase that does not quite fit. Sustained low perplexity across several hundred words is the strongest single tell.

Burstiness: how much sentence rhythm varies

People write a long sentence, then a short one. Then a fragment. Model output is more even, with sentence lengths and structures clustered tightly around a mean. Low burstiness on its own proves nothing, plenty of humans write evenly, but low burstiness plus low perplexity is a much stronger combination than either alone. There is a longer walkthrough in perplexity and burstiness explained.

Sentence-level scoring, not just a document number

GPTOne scores every sentence and returns the highest-scoring ones alongside the document score. This matters more than the headline figure. A document at 64 where three consecutive paragraphs sit above 90 and the rest sit below 20 is a different situation from a document at 64 that is uniformly grey, and only the sentence view tells you which one you are holding.

Reading the result

What each score band means

These are the same thresholds the detection API returns, so what you see on the web and what a developer gets back from an integration agree.

ScoreBandWhat it usually meansSensible next step
0 to 50HumanThe text behaves like human writing across both perplexity and rhythm.Nothing further, unless you have a separate reason to look.
51 to 60MixedOften genuinely mixed authorship, or human text that has been heavily edited for smoothness.Read the sentence breakdown. A cluster tells a different story from an even spread.
61 to 100AIThe text sits well inside the machine-written range.Still a signal, not a finding. Look at which sentences drove it and ask the writer.

Every band is probabilistic. A high score on a 60-word paragraph is far weaker evidence than the same score on 2,000 words, because short texts do not give the classifier enough to work with.

Coverage

Which ChatGPT versions get caught

GPT-3.5 and GPT-4

The most-detected output on the internet. Both models are strongly patterned and most classifiers, ours included, were trained on plenty of their text.

Strong

GPT-4o and GPT-4.5

More conversational and better at varying rhythm, so burstiness alone helps less. The perplexity signal still holds up.

Strong

o1 and o3

The reasoning models write longer and more structured answers. Structure is itself a signal, though it overlaps with careful human academic writing, which is where false positives come from.

Good

GPT-5

49 of 50 samples flagged in our benchmark. Newer models are generally harder because they write with more variation, and this is the number worth rechecking as releases land.

98% in our study

If you need Claude and Gemini in the same pass rather than ChatGPT alone, the multi-model detector page covers all three families. And do AI detectors work on GPT-5 goes deeper on why each new release resets part of the problem.

The limits

Where ChatGPT detection gets unreliable

These are the four conditions that degrade a result most, in rough order of how often they show up.

1. The text was rewritten after generation

This is the big one. On the 50 humanized samples in our benchmark GPTOne caught 39, a 78% rate. That is the best figure in the study and it is still one miss in five. Paraphrasing raises perplexity and breaks the rhythm signature, which is exactly what the classifier reads.

2. The passage is short

Under roughly 200 words there is not enough text for the statistics to settle. Short-text scores swing hard on small changes, so a high score on a paragraph should carry far less weight than the same score on a full essay.

3. The writer is not a native English speaker

Second-language writing tends toward simpler, more regular constructions, which is the same surface pattern a model produces. GPTOne wrongly flagged 4 of 50 non-native samples, an 8% rate. Better than the 34% and 42% we measured for two other tools on the same set, but still double our overall false positive rate. We wrote about the mechanism in why AI detectors falsely flag non-native English writers.

4. The document mixes human and AI writing

A single document score averages across passages that deserve different verdicts. This is where sentence-level results stop being a nice extra and become the only useful output.

Worth being clear about

A detector score is evidence of a statistical pattern. It is not evidence of what a person did. Anyone acting on a score, in a classroom or a newsroom or a hiring pipeline, needs a second source before they act.

Common questions

Questions people ask about ChatGPT detection

Often, yes, and reliably enough to be worth running. On our own 600-sample benchmark GPTOne flagged 49 of 50 GPT-5 samples correctly, a 98% detection rate on unedited output. The number drops once text is paraphrased or rewritten: on the 50 deliberately humanized samples in the same study the detection rate fell to 78%. So a clean paste from ChatGPT is usually caught. A passage someone rewrote by hand afterwards frequently is not.
The OpenAI family across GPT-3.5, GPT-4, GPT-4o, GPT-4o mini, GPT-4.5, o1, o3 and GPT-5. Detection does not work by recognising a version string, it works on statistical patterns in the writing, so newer releases are usually picked up before a version-specific update ships. Coverage also extends past OpenAI to Claude, Gemini, Grok, DeepSeek and LLaMA. See the multi-model detector page if you need all three of the big families in one pass.
It means the model puts the document well inside the range it associates with machine-written text, not that 71% of the words came from ChatGPT. GPTOne reads a score of 0 to 50 as human, 51 to 60 as mixed, and 61 to 100 as AI. Treat the band as the signal and the sentence-level breakdown as the detail worth reading. We go through this in what a 20% AI score means.
Scanning needs a free account. You get free credits on signup with no card required, and one credit covers one word of analysis. Each scan takes up to 50,000 characters, roughly 8,000 words, and that ceiling applies on every plan including Free. Longer documents get checked in sections rather than in one pass.
No, and you should not try. Every detector on the market, ours included, produces false positives. In our benchmark GPTOne flagged 9 of 250 human-written samples as AI, and 4 of the 50 written by non-native English speakers. Those are low rates, not zero rates. A score tells you where to look, and the conversation, the draft history and the writer's own account are what settle it.

Go deeper on ChatGPT detection

Three pieces that cover the parts this page only summarises.

Check a passage against ChatGPT now

Sign up free and your starting credits cover your first scans. Up to 50,000 characters at a time, with sentence-level results rather than one number.

Run a free AI scan

Free credits on signup · No card required · ChatGPT, Claude, Gemini, Grok, DeepSeek and LLaMA