Answer
How much context does an AI agent actually need?
Enough that every decision in the job is covered by something present, and nothing else. The second half is the hard part.
Enough that every decision the job requires is supported by something present, and no more. Find the amount by listing the decisions and asking what each one needs, rather than by gathering what is available.
The question is usually approached from the supply side: here is the material, how much of it should I include. That framing has no stopping rule, because everything is arguably relevant and nothing is obviously excludable. Approached from the demand side it becomes tractable. Write down the decisions the job actually contains — which category, which record, which approach, whether to proceed — and ask what each one needs to be made correctly. The union of those answers is the requirement, and it is normally much smaller than the folder.
This exercise also surfaces the decisions that cannot be supported, which is the more valuable output. If a decision depends on knowing which of two records is authoritative and nothing available says so, then no quantity of context fixes it and the job needs either that fact written down or that decision routed to a person. Discovering this before the run is considerably cheaper than discovering it in the output.
There is a second requirement beyond the decisions: the agent needs enough to know what it is looking at. A record supplied without knowing what the fields mean, a document without knowing whether it is current, a codebase area without knowing what depends on it. These are orientation costs and they are real, but they are bounded and specific, and they are better met by a short note than by supplying more material and hoping the context is inferable.
The reason to be strict about the upper bound is that the cost is not only money. Volume degrades the ability to use a fact that is genuinely present, so an agent given the relevant document inside a large pile can perform worse than one given the document alone. This is counter-intuitive and it is the single most useful thing to know about the subject, because it means the instinct to include more is actively harmful past a threshold rather than merely wasteful.
In practice the method that works is to start deliberately short and add on evidence. Run the job with what the decisions require. Where it goes wrong, look at whether the failure was caused by something missing, and add exactly that. Most teams discover the working set is smaller than they assumed and that two or three specific additions do all the work, which is a far better position than one arrived at by including everything and never learning what mattered.
The right amount of context is decided by the decisions, and a folder is not a list of decisions however relevant it feels.
Siddharth Sharma, Context Theory
Related questions
Should the agent fetch material itself rather than being given it?
Where the job spans more material than the decisions need at once, yes, because holding identifiers and fetching on demand keeps the working set small. The cost is a risk that it does not fetch something it needed, which shows up as a confident answer built on an incomplete look. That risk is manageable when the completion condition requires stating what was read.
Is a summary as good as the source?
For orientation, usually. For any decision that turns on a specific — a value, a wording, an exception — it is not, because summarisation removes exactly the particulars that decisions depend on and does so without announcing which. The workable split is summaries for navigation and the source for anything the answer rests on.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Visibility lift in AI-generated answers from GEO methods | up to 40% | Category-wide |
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
Aggarwal et al., "GEO: Generative Engine Optimization", Princeton / Georgia Tech / IIT Delhi / Allen Institute for AI — KDD 2024 · GEO-bench · 10,000 queries across 8 domains · verified
2026 speed-to-lead benchmark · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Sizing context from the supply side has no stopping rule because relevance is always arguable, while enumerating the decisions in the job and asking what each requires produces a bounded requirement that is normally far smaller than the available material. | Listing the decisions in a candidate job and the specific input each needs, then comparing the union against the material currently supplied. |
| Response | The decision inventory surfaces requirements that no available material can satisfy, which identifies before the run the choices that must be written down or routed to a person rather than discovered as errors in the output. | Marking which decisions in the inventory have no supporting input among the assembled material. |
| Software | Supplying a needed document inside a large volume can produce worse results than supplying it alone, because retrieval of a present fact degrades with total volume, which makes over-inclusion harmful rather than merely wasteful past a threshold. | Running the same task with the relevant document alone and embedded in a large corpus, and comparing use of a specific fact from it. |
| Buying behaviour | Starting short and adding only what a demonstrated failure required produces a smaller working set and identifies which additions carry the value, whereas starting complete never reveals which material mattered. | Recording each addition made in response to a failure and testing whether removing it reproduces the failure. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one