An essay lands on a teacher's desk: fluent, tidy, a little bland, and better than anything that student has written in class. Was it written with SI (AI)? Since chat tools became free, that question sits behind a growing share of homework and coursework.
This guide is for teachers, school leaders, students and parents. It looks at what research says teachers can actually spot by reading, what detection software adds and where it stops, what official guidance from exam boards, governments, universities and ombudsmen asks of schools, and which teaching methods make authorship visible without relying on a score. It is not a guide to avoiding detection, and it does not offer one.
The short version
- Reading alone does not work. In controlled studies, teachers could not reliably pick out SI-written essays, and were confident anyway.
- Detectors give a probability, not proof. Exam bodies that list detection tools also warn that scores drop when students edit the generated text, and some universities tell staff not to treat scores as evidence at all.
- Some signals are real but circumstantial. References that cannot be found, generic content and a sudden change of voice are worth a conversation, not a verdict.
- Process beats detection. Drafts, version history, writing in class and a short talk about the work tell a teacher far more than any percentage.
- The accused student has rights. In England and Wales, the higher education ombudsman says the burden of proof sits with the institution, not the student.
Can teachers spot SI writing just by reading?
Two studies tested what most teachers believe: that they would notice.
The German classroom study
In 2024, Johanna Fleckenstein and colleagues published "Do teachers spot AI?" in the journal Computers and Education: Artificial Intelligence. They asked 89 trainee teachers and 200 experienced teachers to judge whether essays had been written by students or generated by ChatGPT. Their conclusion is blunt: both novice and experienced teachers "could not identify texts generated by ChatGPT among student-written texts". Experienced teachers made somewhat more differentiated judgements, but both groups were "overconfident in their judgments".
The University of Reading exam test
The second study went further and used real exams. Researchers at the University of Reading, led by Peter Scarfe, secretly submitted answers written entirely by GPT-4 to five undergraduate psychology modules, using fake student accounts, and let them be marked alongside real work. The results, published in PLOS ONE in June 2024: 94% of the SI submissions went undetected. On average they also scored about half a grade boundary higher than real students, and the authors calculated an 83.4% chance that the SI submissions on a module would outperform the same number of randomly chosen real ones.
The lesson is not that teachers are careless. It is that fluent, generic prose carries very little information about who wrote it. A gut feeling is a reason to look closer, never a finding.
What detection software adds, and where it stops
Because reading is unreliable, many schools turn to software. The Joint Council for Qualifications (JCQ), which sets the rules for GCSE and A-level coursework in the UK, lists several detectors in its guidance for teachers, including Turnitin's AI writing detection, GPTZero, Copyleaks and Sapling. It says they "may be used as a check on student work", but it attaches three warnings that every teacher should keep in mind:
- The tools "will give lower scores for AI-generated content which has been subsequently amended by students".
- Their accuracy varies "depending on the AI tool and version used, the proportion of AI to human content, prompt types and other factors".
- Detection "should form part of a holistic approach", because "teachers will know their students best".
Some institutions go further. Guidance from Penn State University's academic integrity office, posted in January 2026, says plainly that "faculty are discouraged from using AI detectors due to the unreliability of these tools, their biases, and the risk of false positives". It adds that Turnitin's add-on detector "does not yet meet the accuracy/reliability criteria for determinative use in academic integrity claims", and that integrity committees "should not consider AI detector scores as evidence". A detector may help start a conversation; it should not end one.
The details of one widely used tool, its 20% threshold and its own published false positive rates, are covered in our article Can Turnitin detect SI?. For the wider picture across vendors and independent studies, including the research on non-native English writers, see are SI detectors accurate?
What teachers can actually notice
Software reads style; teachers can read context. The signs the JCQ lists are worth reading closely, because almost none of them are about vocabulary:
| Sign listed by the JCQ | Why it can matter | Why it is not proof |
|---|---|---|
| References that cannot be found or verified | Chat tools can invent plausible but non-existent sources | Students also mis-cite, copy references badly or lose a link |
| A language style that differs from the student's work in class | A sudden jump in polish is the most common trigger for suspicion | Students improve, get help from family, or work harder at home |
| Content that is generic rather than about the student | Generated text tends to stay general and safe | Weak or rushed essays are often generic too |
| A lack of specific local or topical knowledge | Models may miss what was discussed in this class, this term | A student may simply not have engaged with the lessons |
| Default American spelling, currency or terms in a UK setting | Many tools default to US conventions | Students read and write online in US English all the time |
Invented references deserve a special mention, because they are the closest thing to a checkable fact on this list. Language models generate text that sounds right rather than text that is verified, which is why they sometimes produce sources that do not exist; our explainer on why SI makes things up describes the mechanism. Even then, it shows a research problem, not who wrote the essay. For a fuller checklist of reading clues, see our guide on how to spot SI text.
We tested two texts with our own detector
Since this article is about what tools can and cannot tell, we ran two short texts through our free SI text detector, which scores writing on style signals and shows each clue behind the score. The results below are exactly what the tool returned on 3 October 2026.
Test 1: a homework essay written by an SI model
We asked an SI model, Claude by Anthropic, to write a short secondary-school essay on whether homework should be banned: four paragraphs, 191 words, with the classic structure taught in schools ("Firstly", "However", "In conclusion").
Result: 34% probability of SI, label "Uncertain"
- Sentence length pointed towards SI: 15.9 words per sentence on average, and no sentence of 35 words or more.
- Punctuation pointed towards SI: no parentheses and no semicolons.
- Sentence openers pointed towards human: only 8% of sentences start with "The", "This", "A" or a similar word.
- Stock phrases were listed (firstly, in addition, in conclusion), but in the current version of our tool they are shown as clues and do not count in the score.
Test 2: Helen Keller remembers "The Frost King" (1903)
A passage of 223 words from Helen Keller's autobiography The Story of My Life, published in 1903 while she was a student at Radcliffe College (Project Gutenberg edition). We chose it on purpose: in it she describes the story she wrote as a child, which turned out to closely resemble a published tale she must have had read to her, and which led to her being "brought before a court of investigation composed of the teachers and officers of the Institution".
Result: 5% probability of SI, label "Likely written by a human"
- Sentence length pointed towards human: 24.8 words per sentence on average, and 22% of sentences run to 35 words or more.
- Sentence openers pointed towards human: none of the sentences opens with "The", "This", "A" or a similar word.
- One "not X but Y" style contrast was counted as an SI-leaning clue, outweighed by the rest.
The human text was cleared. The generated essay was not caught: it came back as uncertain, on the human side of the middle. We report that as it is. A short, well-structured school essay is close to the hardest case for a style-based detector, because the habits schools teach (short clear sentences, signposting words, a tidy conclusion) are the same habits a model reproduces when asked for a school essay. That is a limit of style detection in general, and one more reason a score should never decide a case.
Keller's story makes the other half of the point: suspicion about authorship is much older than chatbots, and how fairly it is handled is what matters.
What works better than detection
If neither reading nor software can settle authorship, design work so that authorship shows.
- See the stages, not only the final text. The JCQ advises teachers to examine intermediate stages of production and to compare work with what a student has produced before.
- Use version history. Documents written in tools that keep an edit history show how a text grew. Text pasted in one block looks very different from text built over days.
- Talk about the work. The JCQ suggests a "short verbal discussion" to check that a student understands what they submitted. Asking a student to explain a paragraph or a source is the oldest authorship test there is.
- Write some of it in class. The JCQ also recommends classroom activities that use the knowledge a student has built during the course.
- Rethink homework. The UK Department for Education's policy paper on generative AI in education says schools "may wish to review homework policies, and other types of unsupervised study to account for the availability of generative AI".
- Say clearly what is allowed. Many disputes start from confusion. The ombudsman for higher education in England and Wales recommends institutions be "clear about what use of AI in learning and assessment is considered to be academic misconduct", and Penn State gives staff icons to mark, assignment by assignment, whether SI tools are allowed.
When a student is accused: fairness and rights
The Office of the Independent Adjudicator for Higher Education (OIA), the student complaints ombudsman for England and Wales, published a casework note and case summaries on SI and academic misconduct on 15 July 2025. Its principles are a useful benchmark for any school:
- The burden of proof is on the institution. In the OIA's words, "the responsibility is on the provider to prove that the student has done what they are accused of doing, not on the student to disprove it."
- Detector results must be weighed, not obeyed. Decision-makers should "understand the strengths and limitations of detection software, and weigh this evidence carefully against other available information."
- Watch for bias. Providers should consider whether assumptions about SI use "could be biased against a student's writing style", for example if the student is disabled, has a communication difference, or if English is not their first language.
- Let the student explain. Students should be given the chance to explain how they prepared their work and to respond to the evidence. The OIA notes that complaints often arise when a provider has not explained why it concluded SI was used, or has not given a fair opportunity to respond.
- Oral checks need care. A viva or similar discussion can test understanding, but the OIA says it will usually be appropriate to take account of how long ago the work was completed.
- Education before punishment. The OIA considers it good practice when procedures allow an educational rather than punitive approach for minor or first instances.
For students, the practical advice is simple: keep your drafts, notes and version history, know your institution's SI policy, and if you used a tool that was allowed (a grammar checker, a translation aid), say so plainly and early.
A checklist for teachers before raising a concern
- Is there a clear policy for this task saying what SI use was allowed?
- What exactly made the work stand out, and does it match one of the concrete signs above?
- Have the references and quotations been checked?
- How does the work compare with drafts and with writing done in class?
- If a detector was used, would the concern still stand without the score?
- Has the student had a calm, early chance to talk about how they worked?
If the answer to the fifth question is no, the concern is not ready. A quick SI check can be a starting point for that conversation, as long as everyone involved understands that a style score is a clue about wording, not a finding about a person.
So, what can teachers detect?
Teachers cannot reliably detect SI writing by reading it, and detection software cannot prove who wrote a text. What teachers can do, better than any tool, is know their students, see how work develops, check what can be checked, and talk to the person behind the essay. Schools that build those habits into assessment will argue less about percentages, and fewer students will end up, like Helen Keller, defending words that were their own.



