Context Theory Get your growth audit

Answer

How should an AI agent handle uncertainty?

By naming what it does not know and what it assumed, in the output, where the reader can act on it.

By stating what it assumed and what it could not establish, in the output rather than in its tone. Hedged prose signals discomfort without telling anyone which part is doubtful.

The default behaviour when a system is uncertain is to soften the language: might, generally, in most cases, it appears that. This is the least useful available response. Hedged prose spreads the doubt evenly across a passage in which most of the content is fine, so the reader either discounts everything or, more commonly, reads past it. Nothing has been communicated except unease.

The alternative is to move uncertainty out of the prose and into a named list. Assumptions made, with what each one would change if wrong. Things that could not be established, with what was tried. Places where two sources disagreed, with both readings. These are checkable statements, and each one points at a specific action a reader can take, which is the entire difference between useful and decorative uncertainty.

The most important distinction inside this is between not knowing and not having looked. An agent that reports it could not determine something has said one of two very different things: the information does not exist in what it can reach, or it did not go and find out. Requiring the report to say what was attempted separates these, and it turns out to be one of the most revealing things you can ask an agent to include, because the second case is far more common than the first.

There is a related failure worth naming: the confident bridge. Faced with a gap, an agent frequently fills it with the most plausible value and continues, and the resulting output contains an invented specific presented in the same register as everything around it. This is not solved by asking for more caution. It is reduced by supplying the material and by requiring that any value not present in the sources be marked as supplied by the agent, which makes the bridge visible instead of preventing it — and visible is what you actually need.

Uncertainty about the task is a separate category and should be handled at a different time. If the brief is ambiguous, the useful moment to say so is before the work, not in a note attached to the finished product. A cheap first turn that restates the job as understood and names the ambiguities costs almost nothing and catches the expensive misreadings while they are still free to fix.

Finally, distinguish uncertainty from failure in the reporting, because they need different responses. A run that could not complete because a system was unavailable is a failure and should be retried or escalated. A run that completed while unable to establish two facts is a success with caveats and should be read. Systems that collapse the two into a single status either alarm people over normal caveats or bury genuine failures under them.

Uncertainty expressed as an adverb is decoration; uncertainty expressed as a named assumption is something a reader can check.

Siddharth Sharma, Context Theory

Related questions

Should an agent give a confidence score?

Only if you have checked what the score means in practice, and most people have not. A number attached to an assertion invites arithmetic that its calibration will not support, and it is worse than a named assumption because it looks quantitative. A list of the three things the answer depends on is more useful and cannot be misread as a probability.

What should an agent do when two sources disagree?

Report both and say what distinguishes them, rather than choosing silently. Which is right usually depends on context the agent does not have — which record is authoritative, which document is current — and a silent choice discards the one piece of information a person could have used. This is the same discipline that applies to any reconciliation of conflicting evidence.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

2026 speed-to-lead benchmark · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowHedged prose distributes doubt evenly across content that is mostly sound, so the reader either discounts the whole passage or reads past the hedge, which means softened language communicates unease without identifying what is doubtful.Asking a reader of a hedged output to point at which specific claims the hedging referred to.
ResponseAn inability to determine something is ambiguous between the information being unavailable and the agent not having looked, and requiring a statement of what was attempted separates the two, with the second case proving far more common in practice.Requiring an attempted-steps line on every could-not-determine report and auditing a sample.
SoftwareThe confident bridge — filling a gap with the most plausible value and continuing in the same register — is reduced by marking any value absent from the sources as agent-supplied, which makes the substitution visible rather than preventing it.Comparing outputs on a task with a deliberately missing input, with and without a requirement to mark supplied values.
ConstraintUncertainty and failure require different responses — retry or escalation against reading and judgement — so a status scheme that collapses them either raises alarms over routine caveats or conceals genuine failures beneath them.Checking whether the run status of a completed-with-caveats result is distinguishable from that of an aborted run.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one