Answer
How do you write instructions an AI agent will actually follow?
Few, specific, and each one checkable. Adherence falls as instruction count rises, and emphasis does not compensate.
Keep them few, make each one specific enough to check, and remove anything that is not currently needed. Adherence falls as the list grows, and emphasis added to a long document does not recover it.
The first thing to accept is that instruction-following is not binary and does not scale. A short set of instructions is followed closely. A long set is followed approximately, with the specific items that get dropped varying by run. This is not a defect to be prompted around; it is a property to design against, and the design response is to keep the list short rather than to make the items louder.
That means editing is the main activity, and the test for keeping a line is whether it is currently doing work. Most instruction documents contain three kinds of content: rules that matter now, rules added in response to a single incident that has not recurred, and general advice that was never actionable. The second and third categories crowd out the first, and removing them measurably improves adherence to what remains.
Specificity matters more than emphasis. Always verify before proceeding is not checkable and will be interpreted as a mood. Run the test suite and paste the failing output before making a second change is checkable, and the difference in compliance between the two is large. The reliable test is whether you could tell from the output that the instruction was followed. If you could not, the agent cannot either.
Placement matters and is easy to get wrong. Instructions in the middle of a long document compete with everything around them; instructions attached to the task they govern are followed better than the same words filed under general principles. Where something is critical, it belongs in the immediate brief rather than only in the standing document, and the small duplication is worth the adherence.
Prohibitions need a reason attached. Do not modify the schema is followed until a situation arises where modifying the schema is the obvious route to the goal, and then it becomes an obstacle to work around. Do not modify the schema because three other systems read it and will break survives that situation, because the reason is present when the conflict arises. This is the same reason recorded decisions need their rationale.
Finally, treat repeated non-compliance as information about the instruction. If the same rule is broken across different runs, the rule is ambiguous, conflicts with another rule, or is asking for something the situation does not permit. Restating it more firmly is the natural response and almost never works. Reading it as an ambiguity report and rewriting it is slower and does.
Every instruction you add is spent from the same budget, and the twentieth one is paid for by the nineteen that came before it.
Siddharth Sharma, Context Theory
Related questions
Do capital letters and words like 'must' help?
Marginally and unreliably, and the effect fades as the number of emphasised items grows, which it always does. Emphasis is a relative signal: if several instructions are emphatic, none is. It is worth reserving for one or two genuine invariants, at which point it is doing real work precisely because it is rare.
Should instructions be positive or negative?
Positive where a correct action exists, because a prohibition leaves the alternative unspecified and the alternative is where the surprise happens. Write the results to the review folder is better than do not write to the live folder, since the second is satisfied by any of a large number of places, only one of which you wanted.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
| Close rate — response under 5 minutes vs over 24 hours | 32% vs 12% | Category-wide |
2026 speed-to-lead benchmark · verified
Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Instruction adherence degrades as the instruction count rises and the items that get dropped vary between runs, so shortening the list improves compliance with what remains while adding emphasis to a long document does not. | Testing compliance with one specific instruction placed in a short document and in the same document extended with additional rules. |
| Response | An instruction is followable in proportion to whether its execution would be visible in the output, so a rule that cannot be confirmed from the result is interpreted as a disposition rather than as a requirement. | Asking, for each rule in a set, what in the output would demonstrate that it was followed. |
| Software | A prohibition without a reason is worked around once the prohibited route becomes the obvious path to the goal, whereas a prohibition carrying its consequence survives the conflict because the reason is present at the moment the conflict arises. | Comparing adherence to a bare prohibition and to the same prohibition with its consequence stated, on a task where the prohibited route is the shortest one. |
| Constraint | Repeated non-compliance across separate runs indicates ambiguity, a conflict with another rule, or a requirement the situation cannot satisfy, none of which is addressed by restating the rule more firmly. | Reading a repeatedly broken rule against the other rules in the same document for direct conflict. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one