Text detection asks whether a person wrote a sentence. Image detection asks something harder: whether a camera was ever involved at all. This guide covers what can and cannot be established about an image, how to read a confidence score without over-reading it, and how to run a verification you could defend.
Most people think of image detection as a yes-or-no question. It is not. There are three meaningfully different answers, and the one in the middle is the one that matters most in practice.
A camera captured a scene that physically existed. This is the baseline case, and it is also where false positives hurt most — wrongly flagging a genuine photograph can cost someone a grade, a job, a byline or a reputation. A detector's value is judged as much by how rarely it misfires here as by how often it catches a fake.
Baseline caseNo camera was ever involved. The image was synthesised from a text prompt by a generative model. This is the most tractable case, because the entire frame comes from the same non-photographic process. It is also the case most people mean when they ask whether an image is "fake" — and increasingly the least interesting one, because outright synthetic images are often obvious from context alone.
Most tractableThe hard case, and the consequential one. A genuine photograph has had something added, removed or replaced — a person edited out, a background swapped, a face changed, an object inserted. Most of the frame is authentic, so any single overall verdict tends to understate the problem. This is precisely why region-level analysis exists: the useful finding is not "this image scores 64%" but "this specific area does not behave like the rest of the frame."
Hardest and most consequentialThe distinction matters because the three cases carry different stakes. A fully generated image is usually a question of disclosure. An altered photograph is usually a question of intent — someone had a real image and chose to change what it showed. In journalism, insurance, HR investigations and academic integrity work, the second is almost always the more serious finding.
"AI-generated" covers several distinct technologies with very different characteristics. Knowing which one you are likely dealing with tells you how much weight to put on any result.
The technology behind essentially all current image generation. Output is now routinely photorealistic, and the gap between generations is measured in months rather than years. This is the family behind the overwhelming majority of AI images in circulation today, across every domain from stock photography to disinformation.
A real photograph with an AI-altered region. Now built directly into mainstream photo editors and phone galleries, which means it is no longer a specialist capability — removing a person from a photo takes one tap. Because most of the frame is authentic, a single overall verdict is the wrong tool here.
The older technology behind the synthetic profile pictures that filled social platforms and fake-account networks. Largely superseded for creative work, but still heavily used in fraud, catfishing and bot networks because it is free, instant and endlessly repeatable.
A still pulled from an AI-generated video is a different problem from a still generated directly. Video is compressed aggressively before anyone ever sees it, so an extracted frame arrives already degraded. As generated video becomes ordinary, this category is growing fast, and it is currently the weakest ground for every image detector on the market.
A number on its own is close to useless, and is the single most misread part of any detection tool. GPTOne returns a confidence score together with the evidence behind it, because the evidence is what you actually act on.
It expresses how strongly the available evidence points one way. It is not a probability that a court would accept, and it is not a percentage of the image that is fake. Two images with identical scores can warrant completely different responses depending on what produced them.
A region-level view shows which areas of the frame drove the result. This is the difference between "something about this image is unusual" and "this specific area does not match the rest of the photograph" — the second is actionable, the first is not.
Each result comes with the supporting evidence laid out, so you can see whether a verdict rests on one strong indicator or several weak ones. A high score built from a single signal deserves more scepticism than a moderate score built from agreement across many.
A downloadable report captures the result, the evidence and the file it applied to. For editorial, academic and HR workflows this matters: a decision that gets challenged later needs a record of what was assessed and when.
These constraints apply to every image detector available today, GPTOne included. A tool that does not tell you about them is not being straight with you.
A screenshot is a new, lower-quality copy of an image rather than the image itself. Detail is lost the moment it is taken, and any assessment made from it is correspondingly weaker. Wherever you can, work from the file as it was originally saved or received.
Messaging apps and social platforms routinely re-compress and resize what you upload. An image that has passed through several hands is not the image that was created, regardless of its origin. Read results on a forwarded copy as materially less certain.
Aggressive noise reduction, beauty filters, modern computational photography and very clean studio lighting all smooth away the irregularities of an ordinary photograph. These are the most common sources of false positives, and every one of them is normal photographic practice rather than deception.
Each generation of image models closes gaps the previous one left open. Detection is a moving target with no finish line, which is why any accuracy figure is a snapshot against the models that existed when it was measured rather than a permanent property of a tool.
Thumbnails, avatars and heavily compressed images simply carry less to assess. Results on them are weaker across every tool on the market, and a confident-looking verdict on a tiny image should be treated with suspicion.
No image detector is an appropriate standalone basis for disciplinary action, an employment decision, a publication retraction or a legal claim. A confidence score should open a verification process, not conclude one. Where the outcome carries legal or financial weight, consult a qualified forensic examiner.
The workflow below is what separates a defensible verification from an accusation built on a number someone did not understand.
Request the file as it was originally saved, not a screenshot or a download from a social post. It is the step most often skipped and the one everything else depends on.
Open the breakdown and the heatmap. A result driven by one localised region means something very different from one spread evenly across the frame.
Check the source, the capture context and the chain of custody. Reverse image search. Ask the person who supplied it. Technical evidence is one input among several.
Match your response to the stakes. For anything with legal, academic or financial consequences, a qualified forensic examiner is the appropriate authority, not a web tool.
Related detection guides
Most real verification work involves both. These guides cover the text side in the same depth.
Check a JPG, PNG or WebP up to 10MB. Confidence score, evidence breakdown, region heatmap and a downloadable report.
Check an image →The companion guide for written content — how detection performs across model families, and why the model matters more than the score.
View guide →Model-by-model text detection breakdowns across every GPT, Claude and Gemini release, with accuracy benchmarks.
View guide →Drop in a JPG, PNG or WebP up to 10MB and get a confidence score, an evidence breakdown and a region heatmap in seconds.
Try the free AI Image DetectorFree · No sign-up · Midjourney · DALL·E · Stable Diffusion · Flux · Imagen · Firefly · Sora