Answer
How do you stop an AI tool making things up?
You constrain what it can draw on and you check what it produces. Instructing it to be accurate does not work.
Give it the source material to work from, ask for citations you can check, and verify anything consequential. Instructions to be accurate do not constrain output. Fabrication is most likely where specific facts are requested without being supplied.
The first thing to accept is that instructing a tool to be accurate does not make it accurate. The instruction is understood in the sense that the output becomes more hedged, and the underlying tendency to produce plausible specifics that were never supplied is not addressed by asking. The controls that work operate on what the system has to work with and on what happens to its output, rather than on how it was asked.
The single most effective control is supplying the material. A question answered from a document you provided is a different operation from a question answered from general knowledge, and the failure rate differs accordingly. Paste the policy, the specification, the transcript, the page — then ask about it. The improvement is large and it is available in every tool, including the simplest.
The second is requiring the answer to point at its source in a way you can check. Not a citation in general, but a quotation from the material you supplied, with enough surrounding text to locate it. A claim that cannot be tied back to something checkable is exactly the claim most likely to be invented, and asking for the tie makes that visible rather than preventing it — which is the useful outcome, because you can then discard it.
The third is knowing where fabrication concentrates, because it is not uniform. Specific numbers, dates, names, citations, statutory references, product model numbers, quotations attributed to people — these are where invention happens, because the shape of the answer requires a specific and the specific was not available. General explanation is comparatively safe. So the verification effort should go to every specific in the output, and can largely skip the prose.
The fourth is structural: decide what happens when the system does not know. Tools and configurations differ enormously in whether they decline or continue, and it is testable before you rely on anything — ask about something that does not exist and see whether it says so or produces a confident description. A system that invents a plausible answer to a question about a non-existent thing will do the same for a real question it lacks the material for.
None of this is a reason for avoidance, and it is worth saying so plainly. It is a reason for a working habit: supply the material, ask for the tie back to it, check every specific, and never let an unchecked specific reach a customer, a regulator or a contract. That habit costs a minute per output and it is the difference between a tool that is useful and one that eventually produces something you have to explain.
Asking a system not to make things up is like asking it to be taller; the request is understood and the constraint is not where the request can reach.
Answer Production Engine, Context Theory
Related questions
Do newer models fix this?
They reduce it and do not remove it, and there is a hazard in the improvement: as outputs become more reliable, the checking habit erodes precisely because it usually finds nothing. The remaining errors are then more likely to reach a customer, because nothing was looking. Treat improved reliability as a reason to check faster rather than as a reason to stop.
Is a tool connected to our own documents safer?
Meaningfully, and with a specific residual failure. Retrieval from your own material removes most invented specifics, but the system can still misread what it retrieved, combine two documents that should not be combined, or answer from general knowledge when the retrieval found nothing. Asking for the source passage catches all three, which is why the citation habit still applies.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| AI-cited sources that also rank in the Google organic top 10 | 10% | Category-wide |
| Visibility lift in AI-generated answers from GEO methods | up to 40% | Category-wide |
| Firms that never responded to a web enquiry at all | 23% | Category-wide |
2026 generative engine citation study · fewer than · verified
Aggarwal et al., "GEO: Generative Engine Optimization", Princeton / Georgia Tech / IIT Delhi / Allen Institute for AI — KDD 2024 · GEO-bench · 10,000 queries across 8 domains · verified
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Software | Instructing a system to be accurate changes the hedging of its output without constraining the production of plausible specifics that were never supplied, so the effective controls act on inputs and on verification instead. | Comparing outputs to the same question with and without an accuracy instruction, checked for invented specifics. |
| Workflow | Fabrication concentrates in specifics — numbers, dates, names, citations, statutory references, model numbers and attributed quotations — because the answer's shape demands a specific the system did not have, while general explanation is comparatively safe. | Marking each specific in a sample of outputs and checking only those against source material. |
| Software | Whether a tool declines or continues when it lacks material is testable before reliance by asking about something that does not exist, and a system that describes a non-existent thing confidently will do the same for a real question it lacks material for. | Asking the configured tool about a plausibly named thing that does not exist and recording the response. |
| Software | Retrieval from a business's own documents removes most invented specifics but leaves misreading, improper combination of two documents, and fallback to general knowledge when retrieval finds nothing, all of which requesting the source passage exposes. | Asking the tool to quote the passage it relied on and locating that passage in the source document. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one