Drop a photo or a paragraph into an SI detector (still called an AI detector almost everywhere) and you get a number in seconds. It looks like a measurement. It is closer to a detective weighing clues, some much stronger than others. This guide explains what those clues are, how SI and AI detectors turn them into a score, and why even the best tool can be wrong. We also ran three texts through our own detector and show the real results, including one where it got it wrong.
The short answer
SI detectors look for three kinds of evidence:
- Statistical traces in the content itself. A model trained on thousands of human and generated examples learns the patterns that separate them: pixel textures in images, word choice and sentence rhythm in text.
- Metadata and provenance records. Camera data (EXIF), synthetic-media labels written by generators (IPTC and XMP), and signed Content Credentials (C2PA) that record where a file came from.
- Invisible watermarks. Hidden signals that some SI providers embed in their output and that only a matching detector can read.
A good detector combines several of these and shows which ones it found.
| Clue | What it looks at | Strength | Main weakness |
|---|---|---|---|
| Classifier | Pixel patterns or writing style | Works on any file, even with no metadata | Probabilistic, fooled by new generators, editing and paraphrasing |
| Metadata and C2PA | EXIF, IPTC/XMP labels, signed provenance manifests | A valid signature from an SI tool is very strong evidence | Screenshots, re-saves and most social networks strip it |
| Invisible watermark | A hidden signal embedded at generation time | Can survive some edits | Only works for providers that add one, and only with their detector |
Clue 1: classifiers that learn from examples
Most SI detectors are, at their core, classifiers. The recipe is simple to describe and hard to do well. You collect a large set of content made by people (camera photos, essays, articles) and a large set of content made by SI models (images from diffusion generators, text from chat assistants). You then train a neural network to tell the two apart. The network is never told what to look for. It finds whatever statistical differences separate the two piles.
For images
Image generators build a picture by gradually removing noise, and that process leaves traces that are hard to see but easy to measure: unusually smooth textures, particular frequency patterns, lighting that is consistent in a way real optics rarely are. A trained classifier picks these up even when the image looks perfect to the eye. Its answer is a probability, for example "0.94 likely generated".
For text
Text classifiers do the same with words: some are neural networks, others use hand-picked features such as word frequencies and sentence length, and many tools combine both.
The catch
A classifier is only as good as its training data. If a new generator appears, or if content is heavily edited, compressed or paraphrased, it can land outside what the model has seen, and the score becomes unreliable.
Clue 2: metadata, labels and Content Credentials
The second family of clues looks at the information stored with a file. For images it is often the most reliable evidence, when it exists.
EXIF camera data
Camera photos usually carry EXIF fields such as make, model and exposure. Generators do not write these by default, so their presence is a mild clue towards a real photo, only mild because EXIF can be copied or faked, and many apps delete it.
IPTC and XMP synthetic-media labels
The IPTC, the body behind the photo metadata standard used by news agencies, defines a "digital source type" vocabulary. One value, trainedAlgorithmicMedia, is labelled "Created using Generative AI" and defined as "digital media created algorithmically using an Artificial Intelligence model trained on captured content". Several SI image tools write this label into the file's XMP metadata. If a detector finds it, the file declares itself as generated.
C2PA Content Credentials
The Coalition for Content Provenance and Authenticity (C2PA) publishes an open standard for recording where a piece of content came from and how it was edited. Its steering committee includes Adobe, Google, Microsoft, OpenAI, the BBC and Sony, among others. The record is called a manifest, which the C2PA explainer describes as a set of provenance statements "that are digitally signed". The manifest also contains a cryptographic hash of the content at the time of signing, so any later change to the file or to the record breaks the match.
When an SI generator signs its output with C2PA and the manifest survives, a detector can read who made the file and verify the signature. That is about as close to proof as SI detection gets. The problem is survival: a screenshot, a re-save in an editor or an upload to many platforms removes the manifest. No manifest does not mean "real". It only means "no record".
Clue 3: invisible watermarks
Some SI providers go a step further and hide a signal in the content itself: a pattern in the pixels of an image, or a statistical bias in the words a model chooses. Google DeepMind's SynthID is the best-known example. The C2PA standard also allows "soft bindings", which use invisible watermarking or fingerprint lookup to find a content record again after metadata has been stripped.
Watermarks are useful but narrow: only content from providers that add them carries one, and only the matching detector can read it.
How SI text detectors work: perplexity and burstiness
- Perplexity measures how predictable a text is to a language model. GPTZero, one of the first public AI text detectors, described it in 2023 as "a measure of how likely an AI model would have chosen the exact same set of words as found in the document". Language models tend to pick likely words, so generated text often has low perplexity.
- Burstiness measures how much that predictability, and the rhythm of the writing, varies across a document. People mix short sentences with long ones and plain passages with surprising turns. Models tend to keep an even pace.
Researchers have pushed the idea further. DetectGPT, a 2023 method from Stanford researchers, asks a language model how the probability of a passage changes when it is slightly rewritten. Generated text, they observed, tends to sit at a local peak of the model's probability, so small rewrites make it less likely, while human text does not behave this way as consistently.
These signals are real, but they are not fingerprints: formal writing, manuals, legal text and second-language writing can all be very predictable. GPTZero itself noted in an update to that 2023 article that it later moved to a deep-learning architecture rather than relying on these statistics alone.
A real test: three texts through our SI text detector
We ran three texts through the SI or Not text detector, with the same code the website uses, and report the output as it came out.
Test 1: a typical SI answer
We asked an SI model (Claude, by Anthropic) to write a short explainer paragraph on composting in its usual assistant style. It produced 150 words, beginning:
"Composting is a simple yet powerful way to reduce household waste and enrich your garden soil. [...] Moreover, compost improves soil structure, helps retain moisture, and supports a thriving ecosystem of beneficial microorganisms. [...] Ultimately, composting is not only an environmentally responsible choice but also a rewarding practice [...]"
Result: 99% SI, "Likely SI-generated"
- Typical SI phrasing: 3.3 per 100 words (moreover, additionally, ensures, ultimately, not only ... but also).
- Sentence rhythm: length variation 0.25, average 21.4 words per sentence, a very even pace.
- Structure: 5 sentences open with a transition word.
Test 2: a human text from 1889
For a text that no SI model could have written, we used the opening of Jerome K. Jerome's Three Men in a Boat (1889), in the public domain via Project Gutenberg: 164 words starting "We were all feeling seedy, and we were getting quite nervous about it."
Result: 1% SI, "Likely written by a human"
- Sentence rhythm: length variation 0.61, average 23.4 words per sentence. Long and short sentences alternate.
- Typical SI phrasing: 0.0 per 100 words.
- Vocabulary: MATTR 0.72, with everyday words repeated ("liver", "out of order") the way people write.
A detail we noticed: the Gutenberg file breaks lines at a fixed width. Pasted as-is, each line was read as a sentence (17 instead of 7) and the rhythm clue flipped, although the result stayed at 1%. The figures above are for normal paragraphs.
Test 3: the same SI model, asked to sound casual
We then asked the same SI model to write about the same topic as a casual first-person blog post, with contractions and an informal tone. It produced 148 words, beginning:
"ok so I finally tried composting this spring and honestly? way easier than I thought. I'd been putting it off for years because I pictured a stinky pile of rotting stuff [...]"
Result: 1% SI, "Likely written by a human". This is wrong: the text was generated.
- Human fingerprints: contractions 2, first person 8, casual words 5, lowercase starts 2.
- Sentence rhythm: length variation 0.72.
- Typical SI phrasing: 0.0 per 100 words.
The third test is the most instructive. Style-based detection catches the default voice of an SI assistant well, but a model asked to write differently simply stops producing the clues. The detector worked as designed and was still wrong.

Why no SI detector is certain
Every detector, free or paid, faces the same limits.
- Paraphrasing and prompting. As our third test shows, a different prompt or a pass through a rewriting tool can remove the stylistic clues.
- Editing and compression. Cropping, filters, screenshots and repeated JPEG compression blur pixel traces and delete metadata.
- New generators. A classifier trained on last year's models may not recognise this year's.
- Bias against some human writers. A 2023 Stanford study published in the journal Patterns tested seven widely used GPT detectors on 91 essays written by non-native English speakers for the TOEFL exam. The detectors misclassified more than half of them as AI-generated, with an average false positive rate of 61.22%, while essays by US eighth-graders were classified almost perfectly.
- Short texts. A few sentences do not carry enough signal. This is why our text detector asks for at least 80 words.
Even the companies that build the models struggle. In January 2023 OpenAI released a classifier for AI-written text, reporting that it correctly flagged 26% of AI-written text while wrongly labelling human text as AI-written 9% of the time. On 20 July 2023 it withdrew the tool, adding a note that it was "no longer available due to its low rate of accuracy".
What this means for you: treat a detector score as one clue among others, like a witness statement, not a verdict. Never use it alone to accuse a student, a colleague or a stranger online.
How SI or Not combines the signals
Our SI image detector
- Our own classifier. The image detector runs our own image classifier on our own server and combines its probability with the metadata clues.
- Metadata reading. It reads EXIF camera fields, scans the file for C2PA Content Credentials (including the name of the tool that signed them) and looks for the IPTC
trainedAlgorithmicMedialabel in XMP. - Combination. A C2PA signature or label from a known SI tool pushes the score to the top of the scale, because it is the strongest evidence available. Complete camera EXIF lowers the score somewhat, but does not override the classifiers, since EXIF can be faked.
Our classifiers were trained mainly on photorealistic images. Cartoons and illustrations that have lost their Content Credentials are a known weak spot, and we say so on every result.
Our SI text detector
The text detector is deliberately transparent. It weighs five families of signals: sentence rhythm (how much sentence length varies), typical SI phrasing (words and turns of phrase that language models overuse), vocabulary diversity (a measure called MATTR), human fingerprints (contractions, first person, casual words, typos, informal punctuation) and structural uniformity (evenly sized paragraphs, lists of three, repeated transition openers). It returns a probability, the clues behind it, and highlights the sentences that look most generated. It does not look for watermarks.
Both tools answer with three bands: likely SI, likely human or real, and uncertain. We do not publish an accuracy percentage, because a figure measured on one test set says little about the next file you upload.
How to use an SI detector sensibly
- Check provenance first: Content Credentials and metadata before any classifier score.
- Read the clues, not just the number. One weak clue is not a signed manifest.
- Use the original file, not a screenshot of a screenshot.
- Give text enough length: a full paragraph or more.
- Look for context: who posted it first, and does a reverse image search find an older source?
- Never treat a score as proof about a person.
If you want to try it on your own content, run a free SI check on our homepage. And if you are wondering where the term SI comes from, our guide explains what SI means.



