Hallucination Hunting: A Protocol for Verifying Output
⚠️ Version note: which models search the web, and how well, changes fast. The protocol below doesn't depend on any particular feature. Last checked: 24 July 2026.
Two lawyers filed a brief full of cases that didn't exist
In 2023 a New York law firm submitted a court filing citing six previous decisions. The cases had names, docket numbers, judges, quotations. They also had a problem: none of them existed. A model had produced them, and the lawyer had asked it whether they were real. It said yes.
The instructive part isn't that a machine invented something. It's that a professional, in his own field, reading text about his own subject, could not tell. The fabricated citations looked exactly like real ones — because producing text that looks right is precisely what the system does.
You already know why from AI-01: the model is generating probable continuations, and a wrong answer comes out of the same process, in the same confident shape, as a right one. There is no tell. Reading more carefully doesn't help.
So you don't detect hallucinations by being careful. You detect them with a procedure you run regardless of how the answer looks.
What you'll have at the end
- A four-step checklist you can run in about three minutes
- A risk map — which categories of claim need checking and which genuinely don't
- One real answer audited, with the failures written down
Prerequisite: AI-01, specifically the part about why this is structural rather than a bug.
Sort the claims by risk (2 min)
Not everything needs verifying, and treating every sentence as suspect is how people give up on the protocol by Wednesday. Almost all fabrication concentrates in a few categories.
| Risk | Claim type | Why |
|---|---|---|
| 🔴 High | Citations, papers, authors, book titles, URLs | The single most fabricated category. Plausible-looking references are trivial to generate |
| 🔴 High | Specific numbers, dates, statistics | A number is a token like any other. "37%" is as easy to produce as "34%" |
| 🔴 High | Law, tax, medical, safety | Varies by country and year, changes often, and being wrong is expensive |
| 🟡 Medium | Named events, quotes, who-said-what | Real people get given words they never said |
| 🟡 Medium | Anything after the knowledge cutoff | It may not know, and may not know that it doesn't |
| 🟢 Low | Explanations of stable concepts | How compound interest works isn't going to be invented |
| 🟢 Low | Code you can run | The compiler is the verifier. It fails loudly |
| 🟢 Low | Structure, drafting, rewriting | There's no fact to get wrong |
The practical rule: verify what you'd have to retract. If a claim ends up in something you submit, publish, send to a person who trusts you, or make a decision on — it gets checked.
✅ Check: take an answer you've received recently and mark every 🔴 claim in it. Most people are surprised how few there are — and how load-bearing they are.
The rest of this walkthrough is for members
Behind this: the full step-by-step, the exercise with a verifiable output, and the downloadable cheatsheet. Everything you've read above stays free, always.