How to spot SI text Signs that writing is SI-generated

Last updated · By the SI or Not editorial team · Sources

Short answer

You cannot tell from one sign, but you can from several together, plus a fact check. SI (AI) text tends to have an even rhythm, stock phrases, tidy structure and no lived detail, and it sometimes invents facts or sources. None of these proves anything on its own: formal human writing shares many of them, and a model asked to sound casual hides most of them. Check the facts, ask for drafts, and use a detector only as one more clue.

SI, for Super Intelligence, is the new US government name for AI (our guide explains what SI means), so "SI text" is simply what most people still call AI-generated text: writing produced by chat models such as ChatGPT, Gemini or Claude. This guide is the text companion of how to spot SI images. It lists the signs worth checking, shows what happened when we ran three real texts through our own detector, and explains why the honest answer is often "maybe".

Signs that a text was SI-generated

Think of these as a checklist. Each sign alone is weak. What matters is how many appear together, and whether the text also fails the fact check in sign 6.

1. An even, metronomic rhythm

People write in bursts: a long sentence, then a short one. Then a fragment. Default chatbot output tends to keep every sentence at a similar length and shape. Detection tools call this variation "burstiness". GPTZero, which popularised the term, describes it as a measure of how much writing patterns vary across a document, and notes that human writing varies more. Read a paragraph aloud: if every sentence lands with the same beat, note it.

2. Stock phrases and favourite words

Language models lean on the same connectors and intensifiers: "moreover", "additionally", "ultimately", "it is important to note", "not only ... but also", "plays a crucial role". This is measurable at scale. Kobak and colleagues analysed more than 15 million biomedical abstracts published on PubMed and tracked the sudden rise of certain style words after chatbots appeared; from that "excess vocabulary" they estimated that at least 13.5% of 2024 abstracts were processed with large language models. The title of their paper, "Delving into LLM-assisted writing", is itself a nod to one of those words.

3. Structure that is too tidy

Evenly sized paragraphs, lists of exactly three items, a neat conclusion that restates the introduction, and sentences that open with a transition one after the other. Human drafts are messier: a paragraph runs long because the writer got interested, another is a single line.

4. General claims, no lived detail

SI text is good at the average version of a topic. It rarely includes the specific, slightly odd detail that only someone who was there would know: the name of the colleague who suggested the idea, the price that surprised them, the mistake they made first. Look for sentences that could be pasted under any headline on the same subject.

5. Hedging and balance everywhere

"It depends on your needs." "There are pros and cons." "Whether you are a beginner or an expert." Chat models are tuned to be helpful and inoffensive, so they often avoid a clear position. A text that never commits to anything is not proof of SI, but it is typical of default outputs.

6. Facts, quotes and sources that do not exist

This is the most concrete sign, because you can verify it. Language models can produce confident, well-formatted references to things that do not exist, a failure usually called hallucination. The best-known case is Mata v. Avianca: in June 2023, Judge P. Kevin Castel of the US District Court for the Southern District of New York fined the plaintiff's lawyers $5,000 after they filed a brief citing fake court decisions generated by ChatGPT. When the lawyers asked the chatbot whether the cases were real, it assured them they "indeed exist". Search for every quote and open every link.

7. A voice that does not match the author

If you know the writer, compare with their earlier work: vocabulary, spelling habits, the way they open an email. A sudden jump to flawless, generic prose is worth a gentle question. The same applies to a support ticket, a review or a cover letter that sounds like none of the others around it.

8. No history behind the document

For a document, the file itself can help. Word processors keep version histories, shared documents show when text was typed or pasted in, and notes or outlines usually exist somewhere. A long essay that appeared in one paste, with no earlier version, is a clue worth asking about, not a conviction.

Our test: three texts through our SI text detector

To see how these signs behave in practice, we ran three new samples through our SI text detector (the analyze_text function behind the tool) and copied the results below exactly as the tool returned them.

Text A: written by an SI model in its default style

We asked an SI model (Claude, by Anthropic) for a short paragraph on learning a second language in a typical assistant style. It returned 141 words, beginning:

"Learning a second language is one of the most rewarding investments you can make in your personal and professional growth. [...] Moreover, research suggests that bilingual individuals often demonstrate stronger problem-solving skills and improved memory. [...] Ultimately, mastering a new language is not only a valuable skill but also a meaningful way to connect with people around the world."

Result: 99% SI, "Likely SI-generated"

  • Sentence rhythm: length variation 0.16, average 20.3 words per sentence, a very even beat.
  • Typical SI phrasing: 3.5 per 100 words (enhances, moreover, additionally, ultimately, not only ... but also).
  • Structure: 3 lists of three and 4 sentences opening with a transition.

Text B: the same SI model, asked to rewrite it casually

We then asked the same SI model to rewrite that paragraph in a casual first-person voice, with contractions and asides. It is still machine-written, 140 words, beginning:

"Learning a second language? Honestly, it's one of the best things I've done, and I'm still bad at it. [...] My advice: set small goals. Tiny ones."

Result: 1% SI, "Likely written by a human". This is wrong: the text was generated.

  • Human fingerprints: 5 contractions, 4 first-person marks, 2 casual words, an ellipsis, an aside in brackets and a lowercase sentence start.
  • Typical SI phrasing: 0.0 per 100 words.
  • Sentence rhythm: length variation 0.54, average 11.8 words per sentence.

Text C: a human text from 1854

For writing no model could have produced, we used the opening of Henry David Thoreau's Walden (1854), in the public domain via Project Gutenberg: 262 words beginning "When I wrote the following pages, or rather the bulk of them, I lived alone, in the woods".

Result: 1% SI, "Likely written by a human"

  • Typical SI phrasing: 0.0 per 100 words.
  • Vocabulary: MATTR 0.78, with everyday words repeated the way people write.
  • Human fingerprints: 27 first-person marks.

A detail worth noting: the tool counted three lists of three in Thoreau's text, and it even contains the word "Moreover". Style signs overlap between human and SI writing; it is the combination that tips the score.

The lesson from Text B matters more than the other two. Style signs describe the default voice of chatbots, and anyone can change that voice with one instruction. A low score means "no SI style found", not "a person wrote this". Our article on how SI detectors work explains why every statistical detector shares this limit.

What these signs do not prove

Style is a clue about a text, never a verdict about a person. Before you act on any of the signs above, keep these limits in mind.

How to tell if text is SI generated: a 5-step method

Here is the order we recommend. It starts with what you can verify and ends with the tool, because a score is the weakest evidence of the lot.

  1. 1

    Read it once as a reader

    Before hunting for clues, ask whether the text says anything specific: a name, a number, a place, an opinion someone could disagree with. Generic text that could be pasted under any headline is the first warning sign.

  2. 2

    Check every fact, quote and source

    Search for the quotes, open the links, look up the studies and court cases. Invented references are the most concrete evidence you can find, and they matter whoever wrote the text.

  3. 3

    Look at the style signs

    Rhythm, stock phrases, tidy lists of three, transitions at the start of every sentence, no human fingerprints. One sign means little; several together are worth noting.

  4. 4

    Ask for the process, not a confession

    Drafts, notes, version history, the sources the author used. A person who wrote a text can usually talk about how they wrote it.

  5. 5

    Run a detector last, and read its clues

    Paste the text into an SI text detector and read why it scored the way it did. Treat the score as one more clue, never as the verdict.

If you only have a minute, steps 2 and 5 together already help: a false reference settles more than any score, and the clues listed by a detector show you which sentences to reread.

How to tell if something is SI (AI) generated

The question "is this SI?" does not stop at text. The same logic, verify first and read signals second, works for other media:

When you want a second opinion, our free SI detector answers is this SI for images and text, with the probability and the clues behind it.

Summary: signs of SI text at a glance

How much each sign tells you
SignWhat to look forWeight
Invented facts or sourcesQuotes, studies or cases you cannot findStrong, and verifiable
No historyNo drafts, notes or version history for a long documentMedium, ask first
Voice mismatchText unlike the author's earlier writingMedium
Even rhythmSentences of similar length and shapeWeak alone
Stock phrasesMoreover, additionally, ultimately, not only ... but alsoWeak alone
Tidy structureLists of three, transitions at every sentenceWeak alone
No lived detailGeneric claims, constant hedgingWeak alone
Detector scoreA probability with cluesOne clue among others

FAQ

How can you tell if text is SI (AI) generated?

Look for a combination of signs rather than a single one: an even sentence rhythm, stock phrases such as "moreover" or "it is important to note", tidy lists of three, general statements with no lived detail, and facts or sources that do not check out. Then ask for drafts or notes and, if you want a second opinion, run the text through an SI text detector and read the clues it lists. No sign and no tool is proof on its own.

What words does SI use too much?

Research on scientific abstracts by Kobak and colleagues found that certain style words became sharply more frequent after large language models became available, which is how they estimated LLM use at scale. In everyday text, transition words such as "moreover", "additionally" and "ultimately", and phrases such as "not only ... but also", are common in default chatbot output. A person can use all of them too.

Can SI-generated text be detected reliably?

Not reliably enough to act on alone. OpenAI withdrew its own AI text classifier in July 2023, citing its low rate of accuracy, and research by Liang and colleagues found that several detectors flagged essays by non-native English writers as generated. Rewriting or prompting a model to sound casual can also lower scores, as our own test on this page shows.

Is formal writing more likely to be flagged as SI?

Yes. Polished, formal or template-like writing, including legal and academic prose, shares the regular rhythm and vocabulary that detectors associate with SI. Writers who learned English as a second language are also at higher risk of false positives. That is why a score should open a conversation, not close one.

How do you tell if something is SI generated in general?

For images, check hands, text, reflections and backgrounds, and read the metadata for Content Credentials. For text, check facts, sources and style. For audio and video, look for the original source and trusted reporting. In every case, the context of where the content first appeared tells you more than any single detail.

Does SI or Not store the text I check?

No. Texts are analysed in memory and not saved. A share link keeps only the score, the clues and at most three short excerpts, for 7 days.

Sources

  1. Kobak, González-Márquez, Horvát and Lause, Delving into LLM-assisted writing in biomedical publications through excess vocabulary · arXiv (published in Science Advances, 2025), June 2024, revised July 2025
  2. Liang, Yuksekgonul, Mao, Wu and Zou, GPT detectors are biased against non-native English writers · arXiv (published in Patterns), April 2023
  3. Perplexity and burstiness: what is it? · GPTZero, 1 March 2023
  4. OpenAI scuttles AI-written text detector over "low rate of accuracy" · TechCrunch, 25 July 2023
  5. Mata v. Avianca, Inc. · Wikipedia, accessed September 2026
  6. Henry David Thoreau, Walden (1854), public domain · Project Gutenberg, accessed September 2026

Published , last updated .