Ask an SI (AI) chatbot a simple question and it will almost always answer, fluently and with confidence. Most of the time the answer is fine. Sometimes it is invented from start to finish: a court case that never existed, a quote nobody said, a date that is off by years. Researchers call this a hallucination. This guide explains what an SI hallucination is, why language models produce them, what they have cost real people and companies, and how to catch one before you repeat it. We also ran a made-up paragraph through our own detector to show what SI detection can and cannot tell you about truth.
What is an SI hallucination?
An SI hallucination, still called an AI hallucination almost everywhere, is a statement produced by a language model that sounds plausible but is false or unsupported. OpenAI, in a research paper published in September 2025, describes the behaviour as models that "guess when uncertain, producing plausible yet incorrect statements". Our own guide to what SI is defines it in its glossary as a confident but false statement, such as an invented quote, source or fact.
Two details matter. First, a hallucination is not a lie in the human sense: the model has no intention to deceive and no internal notion of "true" to betray. Second, the problem is the confidence. A person who is unsure usually hedges. A chat assistant often delivers an invented answer in exactly the same calm, well-formatted tone as a correct one, which is why hallucinations are so easy to believe.
The main types of hallucination
A widely cited survey by Lei Huang and colleagues, first posted in November 2023, sorts hallucinations into two families. Factuality hallucinations conflict with facts about the world. Faithfulness hallucinations ignore what the user asked or the material the user provided.
| Type | What goes wrong | Example |
|---|---|---|
| Factual contradiction | A real-world fact is stated wrongly | The wrong year or the wrong telescope for a famous discovery |
| Factual fabrication | Something unverifiable or non-existent is presented as fact | A court ruling, study or book that was never written |
| Instruction inconsistency | The answer drifts from what the user asked | You ask for a translation and get an answer to the question instead |
| Context inconsistency | The answer contradicts a document the user supplied | A summary that adds a figure the source text never mentions |
| Logical inconsistency | The answer contradicts itself | A calculation whose steps are right but whose final result is not |
The distinction is useful in practice. Factual errors call for checking the world: an encyclopedia, an official register, the original paper. Faithfulness errors call for checking the answer against your own input: did the summary really come from the document you pasted?
Why SI makes things up
It predicts text, not truth
A large language model is trained to predict the next piece of text, over and over, on a huge body of writing. That training teaches it grammar, style and a great many facts, because facts appear in text. But the objective itself rewards what is likely to come next, not what is true. When the model has seen a fact many times, the likely answer and the true answer usually coincide. When it has not, the model still produces the most plausible-looking continuation, and plausible is not the same as correct.
Rare facts are the weak spot
The OpenAI paper, by Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang, makes this concrete with arbitrary facts such as birthdays. Spelling mistakes follow patterns a model can learn; the birthday of a little-known person does not. The authors argue that if 20% of birthday facts appear exactly once in the training data, a base model can be expected to hallucinate on at least 20% of birthday questions.
They tested it on one of the authors. Asked for Kalai's birthday and told to answer only if it knew, a state-of-the-art model gave three different dates in three attempts, "03-07", "15-06" and "01-01", none of them correct. Asked for the title of his PhD dissertation, three popular chatbots each produced a different title, university and year. None matched the real thesis, completed at Carnegie Mellon in 2001.
Tests reward guessing
The paper's second argument is about incentives. Most benchmarks score answers as right or wrong. Under that kind of grading, "I don't know" earns exactly as little as a wrong answer, so a model that always guesses scores better than one that admits uncertainty. The authors compare it to a student facing a multiple-choice exam: leaving a question blank guarantees zero, while a guess might earn a point.
OpenAI illustrated the trade-off with its own models on SimpleQA, a set of short factual questions:
| Metric | gpt-5-thinking-mini | OpenAI o4-mini |
|---|---|---|
| Abstention (no answer given) | 52% | 1% |
| Accuracy (correct answers) | 22% | 24% |
| Error rate (wrong answers) | 26% | 75% |
The older model is slightly more "accurate", yet it is wrong three times as often, because it almost never declines to answer. A leaderboard that ranks only on accuracy would prefer it. The authors' proposed fix is not a new hallucination test but a change to the mainstream scoring rules, so that a confident error costs more than an honest "I don't know".
Other causes
The Huang survey groups the remaining causes under three headings: data (misinformation in the training text, knowledge that stops at a cut-off date), training (limits of pre-training, fine-tuning and human feedback, which can reward answers that please the user rather than careful ones) and inference (the way words are sampled one at a time, overconfidence, and failures in multi-step reasoning).
When hallucinations reached the real world
A launch demo with the wrong telescope (2023)
On 8 February 2023, The Verge reported that a promotional demo of Google's new chatbot Bard contained an error. Asked about new discoveries from the James Webb Space Telescope to tell a nine-year-old, Bard said the telescope "took the very first pictures of a planet outside of our own solar system". Astronomers quickly pointed out that the first image of an exoplanet was taken in 2004. According to the European Southern Observatory, astronomers using its Very Large Telescope in Chile spotted the planet 2M1207b in April 2004 and confirmed it in 2005.
Six court cases that did not exist (2023)
In Mata v. Avianca, a personal injury lawsuit against an airline in New York federal court, the plaintiff's lawyers filed a brief citing court decisions such as Varghese v. China Southern Airlines. The cases did not exist. According to the court's opinion, one of the lawyers had used ChatGPT, "which fabricated the cited cases". On 22 June 2023, Judge P. Kevin Castel imposed a $5,000 penalty on two lawyers and their firm. The judge also wrote that "there is nothing inherently improper about using a reliable artificial intelligence tool for assistance": the failure was submitting the output without checking it.
A chatbot's refund policy (2024)
On 14 February 2024, British Columbia's Civil Resolution Tribunal ruled against Air Canada in Moffatt v. Air Canada. The airline's website chatbot had told a customer he could apply for a bereavement fare after his trip, which was not the airline's policy. As summarised by the American Bar Association, Air Canada argued that the chatbot was a separate legal entity responsible for its own actions. The tribunal rejected this and found the airline responsible for all the information on its website, whether it came from a static page or a chatbot.
How common is it?
Rates depend heavily on the model, the task and the question. For legal questions about real US federal court cases, Stanford researchers led by Matthew Dahl and Daniel E. Ho measured hallucination rates between 58% (ChatGPT 4) and 88% (Llama 2) in a study first posted in January 2024. Those figures describe older models on a demanding task, not every chatbot today, but they show why a fluent legal or medical answer deserves extra care.
Can an SI detector spot a hallucination? Our test
A question we get: if SI detection can tell that a text was generated, can it also tell that the text is wrong? We checked with our SI text detector, using the same code as the website.
Test: one true paragraph, one hallucinated paragraph
We asked an SI model, Claude by Anthropic, to write two paragraphs in the same assistant style about the first image of an exoplanet. The first is accurate and checked against the ESO announcement (2004, Very Large Telescope, 2M1207b, about 200 light years away). The second keeps the same wording but swaps in false details: 1998, the Hubble Space Telescope and an invented star called "Tarsis-7". Both texts are SI-generated; neither is presented as human.
"The first direct image of a planet outside our solar system was captured in 1998, not by a ground observatory but from space. Astronomers used the Wide Field Planetary Camera on the Hubble Space Telescope to spot a faint bluish point of light next to a young red dwarf called Tarsis-7. [...]"
Accurate paragraph (120 words): 98% SI, "Likely SI-generated"
Hallucinated paragraph (118 words): 98% SI, "Likely SI-generated"
- Sentence rhythm: length variation 0.23 for the accurate version and 0.22 for the hallucinated one, about 24 words per sentence in both.
- Typical SI phrasing: 1.7 per 100 words in both ("moreover", "ultimately").
- Human fingerprints: none found in either text.
The two scores are practically identical, and that is the correct behaviour. An SI detector reads style: rhythm, phrasing, vocabulary, the statistical traces explained in our article on how SI detectors work. It has no access to the facts, so it cannot tell a true sentence from an invented one written the same way. The reverse is also true: a human can write a false paragraph and a model can write a correct one. Whether a text came from SI and whether it is true are two separate questions, and you need a different method for each.
Keep in mind: a high SI score does not mean a text is wrong, and a low score does not mean it is right. Use SI detection to ask where a text came from, and fact-checking to ask whether it is true.
How to catch a hallucination before you repeat it
- Treat every specific claim as a lead, not a fact. Names, dates, figures, quotes and citations are where hallucinations hide.
- Open the source. If an answer cites a paper, a ruling or an article, find the original. A reference that cannot be found anywhere is a strong warning sign; that is exactly what happened in Mata v. Avianca.
- Ask twice, differently. The Kalai birthday test shows the pattern: when a model is guessing, repeated or rephrased questions often produce different answers. Consistent answers are not proof, but inconsistent ones are a clue.
- Watch for rare or very recent topics. Little-known people, niche technical details and events after the model's training cut-off are where invented details are most likely.
- Check summaries against the document. For faithfulness errors, compare the claim with the text you provided, line by line if it matters.
- Be strictest where the stakes are high. Legal, medical, financial and safety questions deserve a primary source or a qualified professional, whatever the chatbot says.
What SI developers are doing about it
There is no single fix, but several directions come up across the sources above. Grounding a model in retrieved documents, often called retrieval-augmented generation, gives it real text to rely on, although the Huang survey stresses that retrieval has limits of its own and does not remove hallucinations. Training and evaluation that reward abstention, as the OpenAI paper argues, push models to say "I don't know" more often; OpenAI's figures above show a newer model declining to answer about half of the SimpleQA questions instead of guessing. Showing citations lets users check claims, provided they actually click through.
None of this makes hallucinations disappear. The OpenAI authors present them as a predictable result of how models are trained and graded rather than a mysterious glitch, which also means they will not vanish just because the technology gets a new name, as with the renaming of AI to SI in September 2026.
The bottom line
SI models make things up because they are built to produce plausible text and have long been graded in ways that reward a confident guess over an honest "I don't know". The result can be harmless, embarrassing or expensive, as a lawyer and an airline learned. The defence is the same as for any unverified source: check the specifics, open the originals and never let fluency stand in for evidence. For a wider picture of what today's systems do well and badly, see the section on what SI can and cannot do. And when the question is not "is this true?" but "was this made by SI?", our free SI detection tools for images and text give you a second opinion, with every clue shown.



