You paste an essay into an SI detector (an AI detector, in the older vocabulary) and it answers "98% SI". It feels like a lab result. It is not. Every SI or AI detector on the market, ours included, gets some files wrong, and independent researchers have measured how often, sometimes with uncomfortable results. This article looks at the real accuracy of SI detectors: what the studies found, why false positives happen, which groups of writers are hit hardest, and why no detector can ever be 100% reliable. We also ran our own tools on five files and show every result, including the ones where the tool was wrong.
If you want to know how detectors work under the hood (classifiers, metadata, watermarks), read our companion guide how SI detectors work. Here the question is narrower: when a detector gives you a verdict, how much should you trust it?
The short answer
- SI detectors are useful but never certain. They catch a lot of typical, untouched SI output and miss a lot of edited, paraphrased or unusual content.
- Independent tests are far less flattering than vendor claims. A 2023 academic study of 14 text detectors concluded they were "neither accurate nor reliable".
- False positives are real and unevenly spread. Writing by non-native English speakers was flagged as SI far more often than writing by native speakers in a Stanford study.
- Image detectors have the same problem. A 2026 benchmark found that images from recent generators defeated most open detectors.
- A score is evidence, not proof. Never use it alone to accuse a student, a journalist or anyone else.
What independent studies found
Text detectors: 14 tools, one blunt conclusion
In 2023 a team of eight researchers led by Debora Weber-Wulff tested 12 publicly available detection tools and two commercial systems used in education, Turnitin and PlagiarismCheck. They fed them texts written by people, texts generated by ChatGPT, and generated texts that had been machine-translated or manually edited. Their paper, published in the International Journal for Educational Integrity, is direct: "The available detection tools are neither accurate nor reliable." They also found that the tools had "a main bias towards classifying the output as human-written rather than detecting AI-generated text", and that obfuscation, meaning editing or paraphrasing, "significantly" worsened performance.
That bias is a trade-off: a detector that leans towards "human" accuses less, but it also lets more generated text through.
The detector makers themselves
OpenAI, the company behind ChatGPT, released its own text classifier in January 2023. Its launch page was candid: on a "challenge set" of English texts, the classifier correctly identified 26% of AI-written text as "likely AI-written", while labelling human-written text as AI-written 9% of the time. It was "very unreliable on short texts (below 1,000 characters)" and worked significantly worse outside English. On 20 July 2023 OpenAI withdrew it "due to its low rate of accuracy".
Turnitin, whose AI writing detection is used by many schools, publishes its own error figures. In a June 2023 post it said its document-level false positive rate is below 1%, but only "for documents with 20% or more AI writing", and that its sentence-level false positive rate is "around 4%": about 4 in 100 highlighted sentences may in fact be human. It added that 54% of those wrongly flagged sentences sit right next to real AI writing, which makes mixed documents the hardest case. We look at Turnitin in detail in Can Turnitin detect SI?
Non-native writers pay the price
The most cited fairness study comes from Stanford. Weixin Liang and colleagues ran seven widely used GPT detectors on 91 TOEFL essays written by non-native English speakers and on 88 essays by US eighth-grade students. The detectors were near perfect on the US essays. On the TOEFL essays they misclassified more than half as AI-generated, with an average false positive rate of 61.22%. All seven detectors agreed that 18 of the 91 human essays were AI-written, and 89 of the 91 were flagged by at least one detector. The authors warned against using such tools "in evaluative or educational settings".
The reason is not mysterious. Many text detectors reward vocabulary that is varied and sentences that are unpredictable. Someone writing carefully in a second language tends to use common words and safe, regular structures, exactly the profile of machine text.
Paraphrasing and the theoretical ceiling
A University of Maryland team (Sadasivan, Kumar, Balasubramanian, Wang and Feizi) went further in a paper titled "Can AI-Generated Text be Reliably Detected?". They showed that recursively paraphrasing generated text significantly reduced detection rates across several families of detectors, including watermark-based ones, while keeping the text readable. They also linked the best possible performance of any detector to how different human and machine text really are. In plain words: as SI writing becomes statistically closer to human writing, the ceiling for every detector comes down, whatever the method.
Image detectors: new generators beat old detectors
Images are not spared. A benchmark published in February 2026 by Simiao Ren and colleagues ran 23 pretrained variants of 16 open detection methods, without retraining, on 12 datasets totalling about 2.6 million images from 291 generators. The best detector averaged 75.0% accuracy and the worst 37.5%. Images from recent generators such as Flux Dev, Firefly v4 and Midjourney v7 defeated most detectors, which reached only 18 to 30% average accuracy on them. Rankings also changed a lot from one dataset to another: a detector that tops one benchmark can be mediocre on the next.
| Study or source | What was tested | Key finding |
|---|---|---|
| Weber-Wulff et al., 2023 | 14 text detectors, including Turnitin | "Neither accurate nor reliable"; biased towards "human"; editing makes it worse |
| OpenAI, 2023 | Its own text classifier | 26% of AI text caught, 9% of human text flagged; withdrawn in July 2023 |
| Turnitin, 2023 | Its own AI writing detection | Under 1% document false positives above 20% AI writing; about 4% at sentence level |
| Liang et al., 2023 | 7 detectors, 91 TOEFL essays | 61.22% average false positive rate on non-native writing |
| Sadasivan et al., 2023 | Several detector families | Recursive paraphrasing sharply lowers detection |
| Ren et al., 2026 | 23 open image detectors, 2.6 million images | 18 to 30% average accuracy on Flux Dev, Firefly v4, Midjourney v7 |
Why no SI detector can be 100% accurate
Some limits are built into the problem itself.
- There is no hidden label in the content. A sentence or a pixel does not carry its origin. Without a watermark or signed provenance, a detector can only estimate how typical something looks of SI output, and humans sometimes write or photograph in ways that look typical.
- Generators move faster than detectors. A classifier learns from examples of existing generators. When a new model ships, its output can fall outside what the detector has seen, as the 2026 image benchmark showed.
- Editing blurs the line. A text written by SI and rewritten by a person, or a photo retouched with SI tools, is genuinely mixed. There is no single right answer for the detector to find.
- Files lose their evidence. Screenshots, re-saves, resizing and social networks strip metadata and Content Credentials, and heavy compression wipes the fine traces that image classifiers rely on.
- Every threshold trades one error for another. Lower the bar and you catch more SI but accuse more humans; raise it and you accuse fewer people but miss more SI.
The base rate trap
Even a small false positive rate adds up. Take a purely hypothetical detector that wrongly flags 1% of human texts. A teacher who runs 500 honest essays through it over a year should expect around five false flags. If very few students actually use SI, those five innocent students could make up a large share of everyone the tool flags. That is why a detector's error rate on paper says little about the chance that a specific flagged person is guilty.
We tested our own tools: five files, real results
We do not publish an accuracy figure for our own SI scanner, for the reasons above: a number from one test set says little about the file you check tomorrow. What we can do is show real runs, including the misses. All results below come from our tools on 28 September 2026.
Text: one correct call, one clear miss
Test 1: typical SI text. We asked an SI model, Claude by Anthropic, to write a short, formal paragraph on remote work (105 words). Our SI text detector returned 99% likely SI. Clues: 4.8 typical SI phrases per 100 words ("moreover", "evolving landscape", "robust", "foster", "ultimately"), very even sentence lengths (variation 0.21), two lists of three and two transition openers.
Test 2: casual SI text. We asked the same model to write a messy first-person post about a failed sourdough loaf, with slang and short fragments (103 words). Result: 1% SI, likely human. Clues: strongly varied rhythm (0.79), no typical SI phrases, 5 contractions, 7 first-person words, 4 casual words. This is wrong: the text was written by SI.
Test 3: real human text. A 235-word passage from Charles Darwin's On the Origin of Species (1859, public domain via Project Gutenberg). Result: 9% SI, likely human, with no typical SI phrases and long, varied sentences (39.2 words on average).
Test 2 is the important one. Style-based text detection keys on habits of default SI writing. Ask a model for a different voice and those habits disappear. Our detector's known limits go the other way too: formal writing and the writing of non-native speakers can score high, the same bias the Stanford study measured. That is why our text tool highlights sentences and lists its clues instead of giving a bare verdict.
Image: the same picture, three different answers
For images we took one illustration whose origin we know for certain: the cover of our Turnitin article, generated with ChatGPT, and checked three copies with our SI image detector.
Original PNG (1672 x 941, with its metadata): 99% likely SI. Clues: signed Content Credentials (C2PA) from OpenAI, an IPTC label "trainedAlgorithmicMedia" in the XMP metadata, and our own classifier at 99% SI.
Re-saved JPEG, same size, metadata removed: 99.7% likely SI, this time on the pixels alone (our classifier at 99.7%).
Small copy, 400 x 225 pixels, heavy JPEG compression, no metadata (about what a thumbnail or a quick screenshot looks like): our classifier dropped to 14% SI, which would have read as "likely real". A safeguard for small files without camera data raised the final answer to 35%, uncertain, with a "low resolution" hint.
Same image, same generator, three very different pixel scores. The first two copies were easy. The third lost both its provenance record and most of the fine detail the classifier relies on. We ran the same shrink test on other SI covers of ours: most stayed above 90%, but one more dropped to uncertain. What you test is often not the original file, but a degraded copy of it.
How to use an SI detector responsibly
Detectors are not useless; their output is one clue among several. Here is how to read one.
- Read the score as a probability, not a verdict. A result in the middle band means the tool does not know, and saying so is a feature, not a bug.
- Look at the clues, not just the number. A signed Content Credentials record from an SI tool is strong evidence. A classifier score on a small, compressed image is weak evidence.
- Test the best copy you can find. The original file, at full size, with its metadata, gives any detector its best chance. Reverse image search can lead you to it.
- Be careful with short, formal or second-language text. These are exactly the cases where false positives cluster.
- Look for other evidence. Drafts and version history for an essay, the original source and other angles for a photo, independent confirmation for a video. Our guide on how to spot SI images lists the checks that do not depend on any detector.
- Never treat a score as proof against a person. A score can justify a closer look; it cannot settle a case on its own.
If a detector has flagged your own work and you did write it, ask what else the decision was based on, and bring your evidence: drafts, notes, sources, file history. Studies like the Stanford one are worth mentioning if you write in a second language.
The bottom line
Are SI detectors accurate? Sometimes, on typical, untouched content. Independent research consistently finds serious error rates once content is edited, paraphrased, written in a second language, compressed or made by a generator the detector has not seen. Even detector makers have withdrawn tools or published sizeable error figures. That does not make detection pointless: it catches a lot, and signed provenance, when present, is close to decisive. But an honest detector should show its uncertainty, and an honest user should treat its answer as one clue among many. It is also why we built the SI or Not game: people are not perfect detectors either, and seeing your own score is a good lesson in humility. Deepfake videos raise the same issue with higher stakes, as we explain in what is a deepfake.



