Answer
What can an AI agent not do?
It cannot know what it was not told, notice what it never looked at, or judge whether its work is finished.
It cannot know what nobody told it, see anything it did not look at, or reliably judge whether its own work is finished. It also cannot hold responsibility, which is a fact about your business rather than the technology.
There is a useful division here between limits that are moving and limits that are not. Capability limits move: things that failed a year ago work now, and predictions about them age badly. Structural limits do not move, because they follow from what the system has access to rather than from how good it is. The second category is what should shape a design, and it is short enough to hold in mind.
The first structural limit is unstated context. An agent knows the material it was given and what it can read. The tacit knowledge that makes a business run — that this customer is difficult, that this figure is provisional, that this process has an exception nobody documented — is not accessible, and no amount of capability substitutes for a fact that exists only in someone's head. Where a job depends heavily on that kind of knowledge, an agent will produce work that is defensible and wrong, which is worse than obviously wrong.
The second is unobserved state. An agent sees the result of the tools it called. It does not see the thing it did not check, and it usually does not know that it did not check. This is why agents confidently report success on partially completed work: the part they looked at was fine. Verification steps that force a look at the whole, rather than at the part the agent chose, are the specific control for this, and they have to be imposed from outside because the agent cannot notice its own omission.
The third is judging completion. Deciding that work is finished requires comparing the result against a standard, and where the standard is implicit the comparison collapses into asking whether the output looks like the kind of thing that was wanted. It usually does, because producing that is precisely what the system is good at. This is the mechanism behind both premature stopping and endless continuation, and the fix in both cases is an external completion test rather than a better instruction.
The fourth is not technical at all. An agent cannot hold responsibility. If an output is wrong and reaches a customer, a regulator or a contract, the accountable party is the business, and no configuration changes that. This matters practically because it determines what may be delegated: work whose consequences someone must answer for needs a person who has actually examined it, and an approval given without examination is worse than none, because it manufactures a record of review.
Two things often listed as limits do not belong here. Agents are not incapable of long or complex work, though long work needs structure they do not supply themselves. And they are not incapable of accuracy, though accuracy comes from constrained inputs and external checks rather than from asking for it. Treating those as permanent limits leads businesses to under-use the technology in exactly the places where it is strongest.
An agent's blind spot is not the hard part of the problem; it is the part of the situation that lives in somebody's head and was never written down.
Siddharth Sharma, Context Theory
Related questions
Will better models remove these limits?
They will move the capability limits and not the structural ones. A more capable model still cannot read a fact that was never written down, still cannot see a file it did not open, and still cannot be the party that answers for an outcome. Designs that assume the structural limits will be engineered away tend to fail in the same place as the models improve around them.
How do you design around unstated context?
Write it down, or keep the job away from it. Those are the only two options, and the first is more achievable than it sounds: the exceptions and the local knowledge that matter for a specific recurring job usually fill a page, not a manual. The failure is assuming an agent will infer them from the data, which it will, incorrectly and plausibly.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Firms that never responded to a web enquiry at all | 23% | Category-wide |
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
2026 speed-to-lead benchmark · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Agent limits divide into capability limits that move with model improvement and structural limits that follow from access rather than from ability, and only the second class should shape a system design because the first ages out of any recommendation built on it. | Classifying each limitation claimed for a candidate design by whether it would change if the model were replaced. |
| Software | An agent cannot detect its own omission, because it observes only the results of tools it chose to call, which is why partially completed work is reported as complete and why verification must force examination of the whole rather than of the portion the agent selected. | Giving an agent a task whose inputs include a file it has no reason to open, and checking whether its completion report acknowledges not having read it. |
| Constraint | Accountability for an output cannot be delegated to an agent, so work whose consequences a named party must answer for requires a person who has actually examined it, and an approval recorded without examination is worse than none because it manufactures evidence of review. | Identifying who would answer to the customer, regulator or counterparty if the agent's output were wrong. |
| Response | Completion judgement collapses into plausibility assessment where the standard is implicit, which is the shared mechanism behind both premature stopping and unbounded continuation, so an external completion test rather than a stronger instruction is the corrective for both. | Comparing stopping behaviour on the same task with an implicit standard and with a stated observable end condition. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one