Answer
What does AI do badly that people assume it does well?
Arithmetic, counting, knowing what it does not know, and anything requiring it to notice what is absent.
Arithmetic performed in prose, counting things, judging what it does not know, and noticing what is missing. All four produce fluent confident output, which is why they are assumed to work.
The list is not about difficulty. Genuinely hard reasoning is often handled well, and several very simple operations are handled poorly, which is counter-intuitive and is why the failures surprise people. What the weak areas share is that the output is fluent, so nothing in the response signals that this particular operation is one of the unreliable ones.
Arithmetic in prose is the first. Adding a column, computing a percentage, reconciling figures: these produce a plausible number that was composed rather than calculated. The same system will write a correct formula or a correct piece of code to do the calculation, which is the tell — generating a mechanism is reliable, executing one in text is not. Anywhere a number matters, it should come from something that ran.
Counting is the second and is a specific case people find hardest to believe. How many items are in this list, how many times does this appear, how many records match: these are exactly the operations where a confident wrong answer is most likely, and where a person's instinct that it must be easy is strongest. If a count matters, count it with something else.
Knowing what it does not know is the third. There is no reliable internal signal distinguishing an answer built from solid material from one constructed to fill a gap, so the confidence of a response carries no information about which you have. This is the property behind most of the practices in this subject, and it is not improved by asking, because the assessment is produced by the same process as the answer.
Noticing absence is the fourth and the most consequential in real work. What is missing from this document, what has not been covered, which case is not handled, who was not consulted: these require comparison against a complete set that is not present. A system reviewing what is in front of it will assess what is there, and the thing that is not there leaves no trace to assess. This is why coverage has to be checked separately from quality everywhere in this corpus.
Two things frequently listed here do not belong. Long or complex tasks are handled well when structured, and the failures attributed to complexity are usually failures of context or completion testing. And accuracy in general is not a weakness so much as a property that depends on what was supplied, which is a different statement and points at a different fix.
Everything it does badly, it does badly in complete sentences, which is why the list is not the one people expect.
Siddharth Sharma, Context Theory
Related questions
Why is arithmetic hard when reasoning is not?
Because the operations are different in kind: producing the next plausible token is well suited to constructing an argument and poorly suited to executing a calculation, where only one answer is correct and plausibility is no guide. This is also why routing the arithmetic to code works so completely — it changes the operation rather than improving it.
Do these improve with better models?
Arithmetic and counting improve, particularly where the system routes them to a calculation tool, which is a design change rather than a capability one. The absence of a signal about what it does not know and the inability to notice absence are structural, and designs that assume those will be engineered away tend to fail in the same place regardless of model.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Close rate — response under 5 minutes vs over 24 hours | 32% vs 12% | Category-wide |
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified
2026 speed-to-lead benchmark · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Software | The weak areas are not the difficult ones: complex reasoning is often handled well while simple counting is not, and what the failures share is fluent output that signals nothing about the reliability of that particular operation. | Asking for a count of items in a supplied list and comparing against the actual count. |
| Workflow | Generating a mechanism is reliable while executing one in text is not, which is why the same system writes a correct formula and misadds a column, and why any number that matters should come from something that ran. | Requesting both a computed total in prose and a formula for it, then running the formula. |
| Response | There is no internal signal distinguishing an answer built from solid material from one filling a gap, so the confidence of a response carries no information about which case applies and asking does not improve it. | Asking about something that does not exist and comparing the confidence of the answer with a well-supported one. |
| Constraint | Detecting absence requires comparison against a complete set that is not present, so a review assesses what it was shown and the omitted item leaves no trace, which is why coverage must be checked separately from quality. | Submitting a document with a known omission and observing whether a review identifies it. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one