Answer
When should an AI agent ask for human input?
When the choice is not recoverable, when the information is missing rather than uncertain, and when the answer commits somebody else.
Three conditions: the action cannot be undone, the information needed does not exist anywhere it can reach, or the decision commits someone else. Escalating merely because it is unsure produces interruptions nobody reads.
There is a wrong version of this question that produces bad systems: how confident should an agent be before proceeding. Confidence is a poor trigger, because the reported confidence of these systems is not well calibrated against correctness, and because a question raised from uncertainty usually arrives without the material a person would need to resolve it. The person then approves, having learned nothing, and a record of human oversight is created where none occurred.
The better triggers are properties of the situation. First, irreversibility: if the action cannot be undone, the decision belongs to whoever can answer for it, regardless of how confident anyone is. This is the trigger that should be enforced structurally rather than left to the agent's judgement, since an agent deciding whether something is reversible is making exactly the judgement it lacks the context for.
Second, missing information rather than uncertainty. There is a real difference between a system that is unsure between two readings of a document and one that needs a fact nobody wrote down: the customer's actual intent, which of two conflicting records is right, whether an exception applies. The first is a reasoning problem and interrupting adds nothing. The second cannot be resolved by any amount of further work, and continuing means guessing. An agent that can tell these apart escalates rarely and usefully.
Third, external commitment. Anything that binds the business to a third party — a price, a date, a scope, an assurance — is a decision the business makes, not one the work produces. This holds even when the answer seems obvious, because the obviousness is being assessed by the party that would be committing you.
The form of the escalation matters as much as the trigger. A question that says it is unsure and asks how to proceed is nearly useless. A question that states what it found, what the options are, what each implies, and what it recommends can be answered in seconds, and the recommendation is the part that makes it cheap: agreeing takes a moment and disagreeing is where the person's knowledge actually enters. Escalations without a recommendation are how supervision becomes work.
Finally, escalation should not be free for the agent to spend. If interruption costs nothing, the safest-seeming policy is to ask often, and frequent asking degrades the whole arrangement: people stop reading and start approving. Making the agent bundle its questions, or hold them until a checkpoint, keeps each one worth the attention it consumes — which is the resource the escalation is actually spending.
An agent that asks about everything is not being careful; it is transferring its uncertainty to somebody who now has less information than it does.
Siddharth Sharma, Context Theory
Related questions
Should an agent stop and wait, or continue and flag?
Continue on everything the question does not block, then stop with the question and the completed work presented together. That converts an interruption into a review of a mostly finished job, which is both cheaper to answer and better informed. Blocking the whole run on one question is only correct when the answer changes work that has not been done yet.
How do you stop an agent asking about things it should decide?
State the decisions it owns, explicitly and by name. Formatting, ordering, wording, which of two equivalent approaches, how to handle a case the brief already covers. Agents over-escalate when the brief leaves ownership ambiguous, and the fix is the same sentence that would resolve it for a person: these are yours, these are mine.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
| Odds of qualifying a lead — replying within the first hour vs after it | 7× | Category-wide |
2026 speed-to-lead benchmark · verified
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Reported confidence is a poor escalation trigger because it is not reliably calibrated to correctness and because an uncertainty-driven question typically reaches the person without the material needed to resolve it, producing approval without oversight. | Reviewing a sample of confidence-triggered escalations for whether the reader had enough information to answer differently. |
| Constraint | Irreversibility should be enforced structurally rather than left to the agent's assessment, since determining whether an action can be undone requires exactly the situational context the agent lacks. | Checking whether the irreversible actions in a workflow are gated by permission or by the agent's own judgement. |
| Response | Uncertainty between readings and absence of a required fact are different states: the first is a reasoning problem that interruption does not improve, while the second cannot be resolved by further work and continuing means guessing. | Classifying an agent's escalations by whether the needed information existed anywhere it could reach. |
| Buying behaviour | Escalation consumes attention, so a policy where interruption is free degrades into frequent asking and unread approval, which makes bundling questions to a checkpoint a control on the resource actually being spent. | Measuring the proportion of escalations approved without a change over a period of frequent interruption. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one