You are reading Chapter 8 of 36 in the AI Bootcamp, AI Central's free course for professionals who want to use AI in their own work. The chapters build on each other, so if this is your first one, Chapter 1 is the place to start. All of them sit on the AI Bootcamp page, and three more arrive every week.
Catching AI hallucinations takes about 30 seconds a fact: copy the exact specific out of the answer, paste it into a search box, and see whether it exists anywhere outside your chat window. The hard part is not the checking. It is knowing which sentences are worth checking, because a wrong answer arrives in the same clean paragraphs as a right one.
This guide explains why an assistant invents, names the 4 shapes the invention takes, shows why asking it whether it is sure does not work, and gives you a check tiered by what a mistake would cost.
We publish AI systems like this for 300,000+ senior professionals at AI Central, and more of them live in the AI Central Library.
Why a wrong answer does not arrive looking wrong
Most people read an AI answer for plausibility. It reads well, the structure is right, the tone is steady, so it goes into the document. Nothing in the prose marks the sentence the assistant made up, because the fluency is a property of how the text was written, not of how well grounded it is.
The other common response is to distrust all of it and check everything. That advice costs unlimited time to follow, so it gets followed for a fortnight and then quietly dropped, usually just before the week it was needed.
This chapter needs one real piece of output an assistant has produced for you, ideally something built from your own documents. If you do not have one yet, Chapter 7 produces it. Come back here once you have.
It is not lying, and it cannot tell
The assistant is not looking anything up. It is producing the words that most plausibly follow your question, given everything it read while it was being built.
Where a fact turned up thousands of times in that reading, the plausible next words are the true ones, which is why general knowledge holds up. Where a fact turned up once, or never, there is nothing to reproduce. So it produces something with the right shape instead: a date shaped like a date, a citation shaped like a citation, a case name shaped like a case name.
That is where the false statement comes from. It does not explain why the assistant does not simply say so, and the second half is the more useful one. Researchers at OpenAI set it out in a 2025 paper: these systems are trained and scored like students sitting an exam where a blank answer earns zero and a guess sometimes earns a point. Under that scoring, guessing beats admitting ignorance every time.
Which is why the confident tone tells you nothing. A human expert hedges unevenly, firm here and careful there, and you use that carefulness without noticing. An assistant reads the same all the way through, on solid ground and on nothing.
It has improved, and it is not solved. One public benchmark runs the easiest possible version of the test: hand the model a passage and ask it to summarise only what is in front of it. The best models still put something in the summary that the passage did not say, in a few percent of cases. That is the floor, on the simplest task, with the source supplied.
Invention comes in 4 shapes
Each of these has shown up in public in the last 18 months, in work produced by professionals with more to lose than you, and each has a different tell.
1. Invented sources
Citations, links and studies formatted perfectly, that do not exist. A researcher at HEC Paris maintains an open database of court decisions worldwide in which someone filed AI-fabricated material and a judge responded to it. It passed 1,600 entries by mid-2026, more than a thousand of them in the United States, and in a large share of them the person responsible was a practising lawyer. In March 2026 a US federal appeals court fined two attorneys $15,000 each, plus costs, over a brief containing more than two dozen fake citations.
2. Plausible figures
A number of the right size in the right units, that nobody published. Ask for a market size, an adoption rate, a benchmark or a conversion figure, and you get one, in the right order of magnitude, attributed to an institution that plausibly publishes that sort of thing. It looks like research because it is formatted like research. The decimal places are not evidence.
3. Stale certainty
Confident answers about events after it stopped reading. Assistants now search the web for many questions, which helps, and it is not the fix people take it for. In 2025 the European Broadcasting Union and the BBC put more than 3,000 news answers from four major assistants in front of professional journalists across 18 countries. Of those answers, 45% contained at least one significant problem, and the largest single category was not invented facts but sourcing: attributions that were missing, wrong, or did not support the claim they were attached to.
4. Fabricated quotes
Real person, real document, words they never wrote. This is the hardest of the four to catch, because everything around it is real. In late 2025 Deloitte agreed to refund part of an A$440,000 report written for an Australian government department, after academics found fabricated references in it, along with a quotation attributed to a real Federal Court judge, in a real case, pinned to paragraphs that do not exist in the judgment. In June 2026 KPMG withdrew a report after only 5 of its 45 citations checked out.
All 4 shapes have the same thing in common. Every one of them is invisible from inside the answer and obvious from outside it, which is what the check further down is built on.
Asking it whether it is sure is not a check
It is the first thing almost everybody tries. You are asking the thing whose reliability is in question to grade itself, using exactly the process that produced the answer in the first place.
Nothing new enters the conversation. It re-reads what it already wrote and produces the words that most plausibly follow a challenge, which is not the same activity as checking. In a published experiment across ten models and seven tasks, a single follow-up of "are you sure?" made the models change their answer roughly half the time, and accuracy after the challenge was on average lower than before it.
So a retraction is not a correction, and a firm "yes, I am confident" is not evidence of anything. Researchers study this under the name sycophancy, and it is still an open problem in 2026.
What works is asking for something that requires information from outside the conversation. Give me the link. Quote me the exact sentence you took that from. Go and search for it and tell me what comes back. Each of those turns a self-assessment into an artefact you can open yourself, and a link that leads nowhere or a quotation that is not in the document is a real answer.
This is not an argument against ever challenging it. Asking an assistant to argue against its own draft is a genuine technique, and Chapter 21 teaches it properly. Just do not confuse the two jobs. Critique can improve reasoning. Only something from outside settles a fact.
Treat the whole conversation as one source. Questioning it again does not add a second one.
Check what a mistake would cost
Change the question. Not "is this true", which you cannot answer about every sentence you read. Ask "what happens if this is wrong, and who finds out". That takes about a second, and it sorts almost everything you do into three piles.
1. Nobody but you will ever see it
No check. Rewriting your own email in a warmer tone, tightening a paragraph you already wrote, building an agenda, summarising a document you have read: in all of those you already know the content and the assistant is only putting it into words. You will catch a mistake by reading the output, which you were going to do anyway. Checking work in this pile is how people run out of patience and stop checking the work that mattered.
2. It leaves your desk with a fact in it
Check every specific. Every number, name, date, price, statute, statistic, study and quotation. Copy the exact string, paste it into an ordinary search, and see whether it exists anywhere outside your chat window. That is about 30 seconds a fact. If it does not come back, it does not go out.
The cost does not scale with the length of the document. It scales with the number of specifics in it. A one-page memo carrying two figures costs you a minute. A report carrying 30 citations costs a quarter of an hour, and if that feels like too much for the piece, the right response is to use fewer citations, not to skip the check. That quarter of an hour is precisely the step Deloitte and KPMG did not take.
3. Wrong costs money or trust
Open the source. Money moves, a contract gets signed, a regulator reads it, a client relies on it, or it goes public with your name on it. Here, existence is not enough. Open the source and read the sentence that is supposed to support the claim, because the most common failure is not a fake source at all. It is a real one that does not say what the answer said it says.
If you cannot say what being wrong would cost, use the row above. And decide which pile a piece of work sits in before you start writing it, not after. It takes 5 seconds against a blank page, and considerably longer once you have a finished draft you have grown fond of.
The chapter ends with 3 exercises that take about 10 minutes: mark up every specific in something an assistant wrote for you, run one of them through the 30-second check, and name the recurring piece of your work that belongs in the top pile.
What actually changes
Before: you read AI output for whether it sounds right, so you catch the clumsy mistakes and miss the confident ones.
After: you can say in one sentence why a confident answer can be false, you recognise the 4 shapes on sight, and your checking is tied to what a mistake would cost rather than to how suspicious you happen to feel that day.
Two things carry into Chapter 9: your marked-up output, and the recurring piece of work you put in the top pile. Chapter 9 asks a different question about that same work, not whether the answer is true but whether it was safe to ask in the first place. If you want a one-page reference to keep beside you while you check, start with our free AI cheat sheets, sharpen the instructions themselves with the 26 principles of prompt engineering, and the rest of the syllabus lives in the AI Central Library.
Frequently Asked Questions
Why do AI assistants make things up?
It is not looking anything up. It produces the words that most plausibly follow your question, so where a fact turned up once or never in what it read, it produces something of the right shape instead. It was also scored on tests where a blank answer earns zero and a guess sometimes earns a point, so it guesses rather than abstains.
How can you tell if an AI answer is wrong?
Not from the prose. The even, assured tone is how the text is written, not how well grounded it is, so there is no tell inside the answer. Copy each specific out of it, number, name, date, price, quotation or source, and check whether it exists outside your chat window.
Does asking ChatGPT "are you sure?" work?
No. Nothing new enters the conversation, so it re-scores its own words. In a published experiment across ten models and seven tasks, that single follow-up made models change their answer roughly half the time, and accuracy afterwards was on average lower than before. Ask for a link, an exact quotation or a search instead.
Do I need to fact-check everything AI writes?
No, and trying to is why people stop. Anything only you will see needs no check. Anything that leaves your desk with a fact in it needs every specific run through a search box, about 30 seconds each. Anything where being wrong costs money or trust needs the source opened and the supporting sentence read.
What are the 4 types of AI hallucination?
Invented sources: citations, links and studies formatted perfectly that do not exist. Plausible figures: a number of the right size in the right units that nobody published. Stale certainty: confident answers about events after it stopped reading. Fabricated quotes: a real person and a real document, with words they never wrote.






