Short answer: yes, Turnitin can flag text that was likely written by an SI (AI) model, and it says it can also spot SI text that was run through paraphrasing and "humanizer" tools. But its score is a probability, not proof. Turnitin itself states that the result "should not be used as the sole basis for adverse actions against a student", hides scores below 20%, and acknowledges false positives on human writing.
Since 22 September 2026 the US government calls artificial intelligence "Super Intelligence" (our explainer on what SI means covers the renaming), but Turnitin, like most of the education sector, still talks about "AI writing". The technology behind the question has not changed: a model trained to tell generated prose from human prose, running inside a tool that thousands of schools and universities already use for plagiarism checks.
This article sticks to what Turnitin documents publicly, what independent sources have found, and what we saw when we ran real texts through our own detector. It is written for students and teachers alike. It is not a guide to getting past the detector.
What Turnitin says its AI writing detector does
Turnitin adds an "AI writing" indicator to its Similarity Report. According to its guides, the capability is "designed to help educators identify text that might be prepared by a generative AI tool, such as large-language models, chatbots, word spinners, and bypasser tools." The AI percentage is separate from the similarity (plagiarism) score: one does not influence the other, and AI highlights do not appear in the Similarity Report itself.
How the score is produced
Turnitin's FAQ describes the mechanics in some detail. When a paper is submitted, its sentences are extracted and grouped into overlapping segments. Each segment is scored by the model with a value between 0 and 1, and every qualifying sentence inherits the score of the segments it belongs to. Those sentence scores are pooled and then aggregated into the document percentage.
The model itself is described as a deep learning system "based on the transformer architecture", trained on Turnitin's own corpus of academic writing. One detail is worth knowing if you have read other explanations online: Turnitin says its model "is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,'" and that, as a result, "individual predictions may not always be explainable in simple feature-by-feature terms." In other words, you will not get a list of reasons why a paragraph was flagged. If you want to understand what those signals mean in general, our article on how SI detectors work explains classifiers, perplexity and burstiness.
Which SI models it claims to detect
The English-language FAQ lists the models Turnitin says it can detect: several GPT-4o and GPT-5 versions, Gemini models up to Gemini 3.1 Pro preview, Claude models, LLaMA 3.3 and LLaMA 4 Maverick, Mistral Large 3, DeepSeek v3.2, Nova 2 Lite and Grok 4.1, "and tools based on these LLMs as well." New models appear every few months, so check the FAQ for the current list.
Paraphrasers and "humanizers"
Turnitin's release notes show how the detector has been extended over time. In December 2023 it announced that it detects likely SI-generated text "even if it may have been paraphrased using an AI word spinner". In July 2024 the report gained a separate category for text that was "likely AI-generated and then was likely revised using an AI-paraphrasing tool or AI word spinner, such as Quillbot". In August 2025 Turnitin added detection of "the likely use of AI bypasser tools", which it describes as tools that "attempt to modify AI-generated text to appear more human-like". These paraphrasing and bypasser capabilities are, according to the FAQ, "only available for English submissions".
What a submission needs to get a score
| Requirement | What Turnitin states |
|---|---|
| Minimum length | At least 300 words of prose text in a long-form writing format |
| Maximum length | 30,000 words |
| File size | Less than 100 MB |
| File types | .docx, .pdf, .txt, .rtf |
| Languages | English, Spanish, Japanese, Arabic |
| What counts | Prose sentences in paragraphs; not poetry, scripts, code, bullet points, tables or annotated bibliographies |
| Who sees it | Instructors and administrators; the indicator is not visible to students, although instructors can download the report as a PDF and share it |
The percentage covers only the prose Turnitin analyses, so it "is not necessarily the percentage of the entire submission".
How to read the score: the 20% line and the asterisk
A Turnitin AI score can appear in several states. Here are the ones students and teachers ask about most:
- 0%: the model did not identify any qualifying text as likely generated.
- *% (asterisk): the model detected something between 1% and 19%. Since July 2024 Turnitin shows no number and no highlights in this range. Its guide explains why: "Our testing has found that there is a higher incidence of false positives when the percentage is between 0 and 19."
- 20% to 100%: the share of qualifying text the model considers likely generated, or likely generated and then paraphrased.
- No score: the file did not meet the requirements above, the language is not supported, or AI detection was not enabled when the paper was submitted.
An asterisk is not an accusation. Turnitin deliberately withholds scores under 20% because they are the least reliable. Treating "*%" as "a little bit of SI" reads more into it than Turnitin does.
There is also a trade-off in the other direction. Turnitin explains that, to keep false positives low, "there is a chance that we might miss some AI written text in a document", giving the example that a document flagged at 50% "could contain as much as 65% AI writing." The tool is tuned to under-flag rather than over-flag.
False positives: what Turnitin itself acknowledges
Every detector makes mistakes, and Turnitin is more open than many vendors about its own. The key figures, quoted from its documentation:
- Document level: Turnitin aims to keep the false positive rate "under 1% for documents with over 20% of AI writing", which it translates as flagging "a human-written document as AI-written for one out of every 100 fully-human written documents." It says it tests every model update on over 700,000 academic papers written before ChatGPT was released.
- Sentence level: in May 2023, Turnitin's Chief Product Officer wrote that the "sentence-level false positive rate is ~4%", meaning a highlighted sentence may be human-written about 4 times in every 100. The same post reported that 54% of these false positive sentences sat right next to actual SI writing, which is why mixed documents are the hardest case.
- Where false positives cluster: Turnitin reported a higher rate in the first and last sentences of a document, often generic introductions and conclusions, and changed its logic to reduce them. Its FAQ adds that false positives can involve "content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."
- Short documents: with only a few hundred words, the prediction is "mostly 'all or nothing'", so a text that mixes original and generated content "could be flagged as entirely AI-generated."
A rate below 1% sounds small for one class. At the scale of a whole university, it is not.
Known limits and criticism
Universities that switched it off
In August 2023, Vanderbilt University announced it was disabling Turnitin's AI detector. Its reasoning was arithmetic: with around 75,000 papers submitted in 2022, a 1% false positive rate could have meant "around 750 student papers could have been incorrectly labeled as having some of it written by AI." Vanderbilt also cited concerns about bias against non-native English speakers and the lack of detail on how the tool decides. Turnitin's FAQ confirms that administrators can switch the AI feature on or off, so whether you are checked depends on your school.
The non-native speaker question
The bias concern comes largely from a 2023 Stanford study by Weixin Liang and colleagues, "GPT detectors are biased against non-native English writers", later published in the journal Patterns. The researchers ran 91 TOEFL essays written by non-native speakers through seven popular detectors and found an average false positive rate of 61.22%; 18 essays were flagged by all seven, and 89 by at least one. It is important to be precise here: Turnitin was not one of the seven detectors tested (they were Originality.AI, Quil.org, Sapling, OpenAI, Crossplag, GPTZero and ZeroGPT).
Turnitin responded in October 2023 with its own evaluation on writing from English language learners. It reported that, for documents of at least 300 words, the difference in false positive rate between non-native and native writers was "small, and not statistically significant", while documents under 300 words showed a larger gap. Useful evidence, but it is the vendor testing its own product, mostly on short secondary-level essays; independent replication would carry more weight.
Our own test: formal human writing can look like SI
We cannot run texts through Turnitin, which is only available to institutions, but we can show the underlying problem with a detector we control. We pasted three texts into our SI text detector and report the raw output below. Our tool is not Turnitin, uses different methods, and would not give the same numbers; these results only illustrate the general principle.
Test 1: the Declaration of Independence (1776)
The first two paragraphs of the US Declaration of Independence, 229 words, from the National Archives transcription. Written by people, centuries before any language model.
Result: 40% probability of SI, label "Uncertain"
- Sentence rhythm pointed towards SI: length variation 0.28, average 57.3 words per sentence.
- Structure pointed towards SI: 2 lists of three.
- Typical SI phrasing pointed towards human: 0.0 per 100 words.
Test 2: the Gettysburg Address (1863)
Abraham Lincoln's speech, 272 words, from the Bliss copy. Formal too, but with very uneven sentence lengths.
Result: 3% probability of SI, label "Likely written by a human"
- Sentence rhythm pointed towards human: length variation 0.71.
- Typical SI phrasing: 0.0 per 100 words.
Test 3: a paragraph written by an SI model
Three short paragraphs on academic integrity, 161 words, generated by an SI model (Claude, by Anthropic) and deliberately kept in a generic, transition-heavy style.
Result: 99% probability of SI, label "Likely SI-generated"
- Typical SI phrasing: 8.7 per 100 words (crucial, moreover, landscape, it is essential, a variety of, furthermore).
- Structure: 3 lists of three and 6 sentences opening with a transition.
- Sentence rhythm: length variation 0.22.
The lesson is not that our tool, or Turnitin, is broken. It is that stylistic detectors read how a text is written, and some human writing shares those features: long, evenly built sentences, balanced lists, formal register. The Declaration did not reach "likely SI", but it moved well away from "likely human" for reasons that have nothing to do with a machine. Students who write in a very polished, regular style sit closer to that zone than others.
If you are a student and think you were flagged unfairly
Turnitin's documentation is clear that it "does not make a determination of misconduct". The decision belongs to your instructor and your institution's procedures. If a score is raised with you, the most useful thing you can bring is evidence of your writing process:
- Version history. Google Docs and Microsoft Word (with files saved to OneDrive) keep a history of edits. A document that grew over days, with deletions and rewrites, tells a very different story from one pasted in a single block.
- Drafts, notes and outlines. Handwritten notes, earlier drafts, research bookmarks and reading notes all show how the ideas developed.
- Your sources. Being able to explain why you chose a source and where a quote came from is strong evidence of authorship.
- Talk about the content. Offer to discuss the argument or to explain a paragraph in your own words. Knowing your own work well is persuasive.
Stay factual and calm, and read your institution's academic integrity policy. You can point out what Turnitin says about its own limits: the score is not meant to be the sole basis for action.
If you are a teacher looking at a high score
- Start a conversation, not a case. Turnitin describes the score as a way "to start a meaningful and impactful dialogue with their students", and says it should be used alongside "further scrutiny and human judgment".
- Look at the highlights in context. Are they concentrated in a generic introduction or conclusion? Is the paper short, repetitive or mostly lists? Those are the cases Turnitin itself describes as more error-prone.
- Compare with what you know. Earlier work by the same student, in-class writing and drafts are better evidence than any percentage.
- Check your course policy. If some SI use was allowed, a high score may simply reflect permitted assistance that was not disclosed clearly.
- Learn the signs yourself. Our guide on how to spot SI text lists what to look for before reaching for any tool, and our free SI checker shows how a style-based score is built, clue by clue.
- Design for process. Staged assignments with drafts, oral follow-ups and personal reflection make authorship visible without relying on a detector at all.
So, can Turnitin detect SI?
Turnitin can detect a lot of SI-generated writing, including text that has been paraphrased or passed through bypasser tools, in English, Spanish, Japanese and Arabic, as long as the submission is long-form prose of at least 300 words. What it cannot do, by its own account, is prove who wrote a text. It hides its least reliable scores, acknowledges false positives at document and sentence level, and asks educators not to act on the number alone. Read that way, the score is one clue among others: useful to open a discussion, never enough to close one.



