Context Theory Get your growth audit

Answer

What makes an AI answer worth acting on?

That you can check the part the decision rests on, and that being wrong about it is something you could recover from.

Two things: the specific claim your decision depends on can be checked, and being wrong about it is recoverable. Neither the confidence of the answer nor its thoroughness bears on either.

Most guidance on this subject is about assessing the answer as a whole, which is both expensive and the wrong unit. A decision rests on one or two specific claims, and the rest of the answer is context that could be entirely wrong without changing what you do. Identifying the load-bearing claim takes a few seconds and reduces the checking problem by an order of magnitude.

Once identified, the question is whether that claim can be checked at all. A claim with an external referent — a value in a record, a passage in a document, a figure in a published source — is checkable in seconds. A claim that is an interpretation or a projection is not checkable, and an answer whose load-bearing element is uncheckable is a starting point rather than a basis for action, however well argued.

The second half is the recovery question, and it is the one people skip because it feels pessimistic. If this is wrong, what happens, and can it be undone. A wrong answer that leads to a reversible action is cheap to be wrong about, which is why a great deal of AI-assisted work is perfectly safe despite an unremarkable error rate. A wrong answer that leads to a message sent, a payment made or a record deleted is not, and the same error rate is unacceptable there.

Those two questions produce four situations and each has an obvious handling. Checkable and recoverable: act. Checkable and not recoverable: check, then act. Uncheckable and recoverable: act and observe, treating it as an experiment. Uncheckable and not recoverable: do not act on this, and find another basis for the decision. The fourth case is where most of the genuine harm sits and it is identifiable in advance.

What does not belong in this test is how the answer reads. Confidence is produced at no cost and carries no information. Length and thoroughness correlate with effort in human writing and not here. Citations improve things only where you open them, which returns to the first question rather than substituting for it. A well-structured answer with an uncheckable load-bearing claim is exactly as risky as a badly structured one and considerably more persuasive.

One practical habit follows. When an answer arrives, the useful first move is to underline the sentence you are relying on. This locates the risk immediately, tells you whether checking is possible, and frequently reveals that the sentence you are relying on is not one the answer actually established.

Find the one sentence the decision rests on, check that, and ignore the rest; the rest was never the risk.

Siddharth Sharma, Context Theory

Related questions

What if several claims are load-bearing?

Then check each, and notice that a decision resting on four uncheckable claims is much weaker than one resting on a single checkable one, regardless of how thorough the analysis was. Multiple dependencies compound rather than reinforce, which is the opposite of how a well-argued answer feels.

Does it matter whether the answer came from AI?

The test is the same for any source and the base rates differ. A colleague's claim carries a track record and someone to ask; a published figure carries a method you can inspect. What changes with a generated answer is that the usual signals of care and hesitation are absent, so the underlining habit does more work than it would elsewhere.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

2026 speed-to-lead benchmark · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowA decision rests on one or two specific claims while the remainder of an answer is context that could be wrong without changing the action, so identifying the load-bearing claim reduces the checking problem substantially.Marking the sentence a decision depends on in a recent answer and testing whether the decision changes if the rest is removed.
ResponseCheckability and recoverability are independent axes producing four cases, of which uncheckable and unrecoverable is where the genuine harm sits and is identifiable before acting.Classifying a set of past decisions by both axes and checking where the bad outcomes fell.
SoftwareConfidence, length and thoroughness carry no information about correctness in generated output because they are produced at no cost, whereas in human writing they correlate weakly with effort.Comparing the confidence and length of a sample of outputs against their measured accuracy.
ConstraintUnderlining the relied-upon sentence frequently reveals that it was not established by the answer at all, which is a detection method available before any external checking is performed.Locating, within an answer, the support for the sentence a decision depends on.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one