Answer
How do you detect hallucinations at scale?
Check what can be checked automatically — identifiers, citations, totals — and sample the rest. Reading everything finds least.
Automate the checks on anything with an external referent: identifiers, links, quotations, totals, dates. Sample the rest. Reading every output is the most expensive method and finds the least, because fluent prose is the easy part.
At one output a person reads it. At a thousand a day that is not available, and the response people reach for — a lighter read — is the worst of both, costing real time and catching the errors that would have been caught anyway. The workable approach divides outputs by whether they contain something a machine can check.
The automatable class is larger than it first appears. Does this identifier exist in the records. Does this link resolve. Does this quotation appear in the cited source. Does this total match the sum. Is this date within the plausible range. Is this value one of the permitted set. Each is a deterministic check, each fails loudly, and each targets exactly the class where fabrication concentrates, because fabrication concentrates in specifics.
That is the important structural point. Invention happens where the shape of the answer required a particular value that was not available, and particular values are precisely the things with external referents. So the automatable checks and the high-risk content are the same set, which is a fortunate coincidence and the reason this approach works at all.
The unautomatable class is judgement: an interpretation, a recommendation, a characterisation of a situation. These cannot be checked mechanically because there is nothing to compare against, and the correct treatment is sampling. A fixed number of outputs read properly each week, chosen randomly rather than by suspicion, gives you an estimate of the rate and, more usefully, shows what kind of errors are getting through.
Random selection matters more than volume. Sampling the outputs someone flagged measures the flagging rather than the system, and it systematically misses the errors that look fine, which are the ones that matter. A small random sample read carefully is worth more than a large convenience sample skimmed, and it produces a number that can be tracked over time.
The last component is the feedback path from downstream. Corrections made by whoever uses the output, complaints, reconciliation differences, rejected records: all of these are detection signals that already exist and are usually discarded. Capturing them costs a field and converts the people already noticing errors into the most reliable detector in the system, without anyone doing extra work.
Anything with an external referent can be checked by a machine, and anything without one has to be sampled by a person; there is no third category worth building for.
Siddharth Sharma, Context Theory
Related questions
Can a model check another model's output for hallucination?
Usefully for support checking — does this passage back this claim — when it sees only the claim and the passage. Less usefully as a general assessment, because judging whether something is invented without a source to compare against reduces to plausibility, which is the property that made the invention pass in the first place.
What sample size is enough?
Enough that the rate stops moving much between periods, which is fewer outputs than people expect when the rate is not tiny. The purpose is trend detection rather than precision: knowing the rate roughly and knowing when it changes is what supports decisions, and chasing a precise figure is effort spent on the wrong property.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Close rate — response under 5 minutes vs over 24 hours | 32% vs 12% | Category-wide |
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified
2026 speed-to-lead benchmark · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Fabrication concentrates in specifics because invention occurs where the answer's shape demanded an unavailable particular, and specifics are exactly the content with external referents, so the automatable checks and the high-risk content coincide. | Classifying detected fabrications by whether they are values with an external referent or general prose. |
| Software | Deterministic checks on identifiers, links, quotations, totals, dates and permitted value sets fail loudly and run at any volume, which makes them the only detection layer that scales with output count. | Running each check against a batch of outputs and counting the failures surfaced without human reading. |
| Response | Sampling by suspicion measures the flagging process rather than the system and systematically misses errors that look correct, so random selection is what produces a trackable rate. | Comparing the error rate found in a flagged sample against a random sample from the same period. |
| Constraint | Downstream corrections, complaints, reconciliation differences and rejected records are existing detection signals produced without additional effort, and capturing them converts the people already noticing errors into the system's most reliable detector. | Checking whether corrections made to outputs are recorded anywhere or applied silently. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one