Answer
How do you check whether an AI answer is correct?
Check the specifics against a source, not the reasoning against your intuition. Fluent reasoning is the least informative part.
Check every specific — numbers, names, dates, citations, quotations — against a source outside the system. The surrounding explanation is not evidence of accuracy, and asking the model whether it is certain changes its wording rather than its correctness.
Verification has to be targeted or it does not happen. Reading an answer carefully feels like checking it and is not, because the reading is being done against your own knowledge, and the cases that matter are the ones where your knowledge is absent. The efficient method is mechanical: mark every specific in the output, and check only those.
Specifics are where errors concentrate. A number, a date, a proper name, a case citation, a statutory reference, a product version, a quotation attributed to someone, a link. Each of these has the property that the answer's shape demanded a particular value, and if the value was not available it will nonetheless be supplied. General explanation is comparatively safe, and it also happens to be the part that is easiest to read, which is why unaided reading finds so little.
The check has to be external. Asking the same system to verify its own claim usually returns agreement, because the question is answered from the same material that produced the claim. Asking a different system is better and still weak, since both may be drawing on the same circulated source. The check that means something is a primary document, a search that finds the thing named, or a query against your own records.
There is a specific test worth running once on any tool you intend to rely on: ask it about something that does not exist. A plausible-sounding regulation, a report with a name you invented, a feature of a product. A system that declines is telling you something useful about how it behaves when it lacks material. A system that produces a confident description has told you what it will do with every real question it also lacks material for.
Links deserve their own line because they fail in a distinctive way. A cited link may not exist, may exist and not contain the claim, or may have contained it and no longer does. Opening the link and finding the sentence is the whole check, and it is the one most often skipped because the presence of a citation reads as verification already performed.
The last part is knowing what you are checking against. Verifying that an answer matches what is commonly said about a subject establishes that it is conventional, not that it is right. Where the question is one where the common answer is wrong, this kind of checking will confirm the error with every appearance of diligence, which is why the strongest verification is against a primary source rather than against consensus.
Confidence is the one thing a language model produces at no cost, which is exactly why it carries no information about whether the answer is right.
Siddharth Sharma, Context Theory
Related questions
Does asking for sources fix this?
It makes the problem visible rather than removing it, which is still worth a great deal. A claim with a citation you can open is one you can dispatch in seconds; a claim without one is indistinguishable from a claim with a source that would not survive being looked at. Requesting citations is useful precisely because it converts a checking problem into a clicking problem.
How much checking is proportionate?
Set it by consequence, not by suspicion. Anything that reaches a customer, a regulator, a contract or a published page gets every specific checked. Anything that informs your own thinking and will be tested by reality shortly afterwards can go unchecked, because the reality test is the verification. The failure mode to avoid is uniform light checking, which costs real time and catches the errors that were going to be caught anyway.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| AI-cited sources that also rank in the Google organic top 10 | 10% | Category-wide |
| Visibility lift in AI-generated answers from GEO methods | up to 40% | Category-wide |
2026 generative engine citation study · fewer than · verified
Aggarwal et al., "GEO: Generative Engine Optimization", Princeton / Georgia Tech / IIT Delhi / Allen Institute for AI — KDD 2024 · GEO-bench · 10,000 queries across 8 domains · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Errors concentrate in specifics rather than in explanation, because a specific is a value the answer's structure required and will be supplied whether or not it was available, while general prose has no equivalent demand for an unavailable particular. | Marking every specific in a sample of outputs and checking only those against primary material. |
| Software | Self-verification returns agreement at a high rate because the check is answered from the same material that generated the claim, and cross-model checking is only marginally stronger where both draw on the same circulated secondary source. | Asking a system to verify a claim it has just made, then locating the primary document the claim rests on. |
| Response | Asking a configured tool about a plausibly named thing that does not exist establishes its behaviour when material is absent, and a system that describes the non-existent thing confidently will do the same on real questions it lacks material for. | Running the non-existent-subject test on the specific tool and configuration before relying on it. |
| Constraint | Checking an answer against what is commonly published establishes that it is conventional rather than correct, so a question whose common answer is wrong will be confirmed by consensus checking with every appearance of diligence. | Tracing any widely repeated statistic to its originating measurement and comparing the two. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one