For educators

AI detector for teachers: a score is where a conversation starts

GPTOne gives you a document score, the sentences that drove it, and a false positive rate we publish rather than hide. What it does not give you is proof, and a process built on the assumption that it does will eventually hurt a student who did nothing wrong.

Free credits on signup, no card required. Up to 50,000 characters per scan, about 8,000 words, on every plan.

Key takeaways

  • 3.6% of human papers were wrongly flagged in our 600-sample study, 9 of 250. In a 120-student cohort that is roughly four papers.
  • 8.0% for non-native English writers, 4 of 50. More than double the general rate, and the single most important number on this page.
  • 78% of rewritten AI text was caught, 39 of 50. A student who paraphrases has a real chance of passing.
  • Sentence-level results matter more than the document score. A clustered flag and an evenly grey document need different responses.
  • Scores are not evidence. Draft history, version data and a conversation are. Build the process around those.
3.6%False positives, human samples
8.0%False positives, non-native writers
600Samples in our first-party study
50,000Characters per scan, all plans
A defensible process

Four steps between a flag and a decision

Institutions that get into trouble skip step two. The gap between "the tool said 89" and "the student cheated" is where every successful appeal lives.

01

Scan, then read the sentences

Open the sentence-level breakdown before you form a view. Three consecutive paragraphs above 90 with the rest below 20 is a specific finding. A document uniformly at 65 usually is not.

02

Look for process evidence

Version history in Docs or Word, intermediate drafts, outlines, research notes. A paper that appeared in one paste tells you something a score never will. So does one built over nine days.

03

Talk to the student, without an accusation

Ask them to walk you through their argument, why they chose a source, what they cut. A student who wrote it can do this easily. Open with the verdict and you lose the only reliable signal you had.

04

Decide on the whole picture, and write it down

Record what you looked at and why you concluded what you concluded. If the case ever gets appealed, the record of a fair process is what holds, not the screenshot of a percentage.

The number that should worry you

False positives are not evenly distributed

A 3.6% false positive rate sounds small until you notice it is not spread evenly across your students. It concentrates on the ones least able to push back.

In our benchmark the non-native English subset was flagged at 8.0%, more than double the overall rate. The same subset was flagged at 34% by one competing tool and 42% by another. Second-language writing tends toward regular sentence structure and conservative vocabulary because that is what gets taught, and regular structure with conservative vocabulary is precisely what a language model produces. The classifier is not detecting AI in those papers. It is detecting careful, learned English.

Practical consequence

If your cohort includes international students, a threshold that works elsewhere will over-flag them. Either raise the score at which you open a review, or accept that the first stage of the review is a conversation and not a sanction. Doing neither produces a pattern of accusations against one group, which is both unjust and institutionally dangerous.

The peer-reviewed literature says the same thing more forcefully than our own data does. We summarised it in what the research shows on AI detector false positives and in why AI detectors falsely flag non-native English writers.

Design around it

Assignment changes that beat detection outright

Require version history

Ask for a shared Doc link rather than a file. The revision timeline is stronger evidence than any detector output, it costs nothing, and it works on every model including ones nobody has benchmarked yet.

Highest value

Grade the process, not only the artefact

Outline, annotated bibliography, draft, reflection. Each stage is cheap to produce honestly and expensive to fake convincingly across weeks.

Structural

Anchor prompts in class specifics

A question that references Tuesday's seminar argument or a local dataset is answerable by a student who was there and generic in the hands of a model that was not.

Prompt design

Add a short oral component

Two minutes defending a thesis separates authorship faster and more fairly than any percentage. It also gives the student a route to clear themselves.

Fastest resolution

None of this replaces a detector. It changes what the detector is for: triage across a stack of papers, rather than a verdict on one. That is a job it is genuinely good at.

Common questions

What teachers ask us about AI detection

No. A detector measures a statistical pattern in text, it does not observe what a student did. In our own 600-sample benchmark GPTOne wrongly flagged 9 of 250 human-written samples, and 4 of the 50 written by non-native English speakers. Those students exist in your classes. Use the score to decide which papers deserve a conversation, and let the conversation, the draft history and the student's own account decide the outcome.
Higher than any benchmark suggests. Our study measured a 3.6% overall false positive rate and 8.0% on non-native English writing, under controlled conditions with one fixed threshold. A real class adds ESL writers, students who draft in a grammar tool, formulaic lab-report styles and rubrics that reward exactly the even, hedged prose a model produces. Plan your process around the assumption that some flags will be wrong.
Process evidence, not scores. Document version history showing a paper built over days, drafts with real revision, notes and outlines, and a student who can discuss their own argument in detail. A detector score is a reason to start looking for that evidence. On its own it has been successfully challenged, and it should be.
Scanning requires a free account. You get free credits on signup with no card required, and one credit covers one word of analysis. Each scan takes up to 50,000 characters, about 8,000 words, on every plan including Free. A long dissertation gets checked in sections rather than in one pass. Paid plans start at $7.99 a month for a larger monthly allowance.
Yes, and it changes behaviour more than the tool does. A stated policy that names what is allowed, what must be disclosed and what happens after a flag removes the ambiguity students currently exploit, and it protects you when a flag turns out to be wrong. We put a starting point in our classroom AI use policy template.

More for educators

Written for teaching staff and academic integrity teams, not for marketing.

Use it for triage, not for verdicts

Sign up free and your starting credits cover your first scans. Sentence-level results tell you which papers are worth a conversation, which is the only question a detector can honestly answer.

Check a submission free

Free credits on signup · No card required · Up to 50,000 characters per scan on every plan