GPTOne gives you a document score, the sentences that drove it, and a false positive rate we publish rather than hide. What it does not give you is proof, and a process built on the assumption that it does will eventually hurt a student who did nothing wrong.
Institutions that get into trouble skip step two. The gap between "the tool said 89" and "the student cheated" is where every successful appeal lives.
Open the sentence-level breakdown before you form a view. Three consecutive paragraphs above 90 with the rest below 20 is a specific finding. A document uniformly at 65 usually is not.
Version history in Docs or Word, intermediate drafts, outlines, research notes. A paper that appeared in one paste tells you something a score never will. So does one built over nine days.
Ask them to walk you through their argument, why they chose a source, what they cut. A student who wrote it can do this easily. Open with the verdict and you lose the only reliable signal you had.
Record what you looked at and why you concluded what you concluded. If the case ever gets appealed, the record of a fair process is what holds, not the screenshot of a percentage.
A 3.6% false positive rate sounds small until you notice it is not spread evenly across your students. It concentrates on the ones least able to push back.
In our benchmark the non-native English subset was flagged at 8.0%, more than double the overall rate. The same subset was flagged at 34% by one competing tool and 42% by another. Second-language writing tends toward regular sentence structure and conservative vocabulary because that is what gets taught, and regular structure with conservative vocabulary is precisely what a language model produces. The classifier is not detecting AI in those papers. It is detecting careful, learned English.
If your cohort includes international students, a threshold that works elsewhere will over-flag them. Either raise the score at which you open a review, or accept that the first stage of the review is a conversation and not a sanction. Doing neither produces a pattern of accusations against one group, which is both unjust and institutionally dangerous.
The peer-reviewed literature says the same thing more forcefully than our own data does. We summarised it in what the research shows on AI detector false positives and in why AI detectors falsely flag non-native English writers.
Ask for a shared Doc link rather than a file. The revision timeline is stronger evidence than any detector output, it costs nothing, and it works on every model including ones nobody has benchmarked yet.
Highest valueOutline, annotated bibliography, draft, reflection. Each stage is cheap to produce honestly and expensive to fake convincingly across weeks.
StructuralA question that references Tuesday's seminar argument or a local dataset is answerable by a student who was there and generic in the hands of a model that was not.
Prompt designTwo minutes defending a thesis separates authorship faster and more fairly than any percentage. It also gives the student a route to clear themselves.
Fastest resolutionNone of this replaces a detector. It changes what the detector is for: triage across a stack of papers, rather than a verdict on one. That is a job it is genuinely good at.
Written for teaching staff and academic integrity teams, not for marketing.
Sign up free and your starting credits cover your first scans. Sentence-level results tell you which papers are worth a conversation, which is the only question a detector can honestly answer.
Check a submission freeFree credits on signup · No card required · Up to 50,000 characters per scan on every plan