Context Theory Get your growth audit

Answer

How do you write instructions an AI agent will actually follow?

Few, specific, and each one checkable. Adherence falls as instruction count rises, and emphasis does not compensate.

Keep them few, make each one specific enough to check, and remove anything that is not currently needed. Adherence falls as the list grows, and emphasis added to a long document does not recover it.

The first thing to accept is that instruction-following is not binary and does not scale. A short set of instructions is followed closely. A long set is followed approximately, with the specific items that get dropped varying by run. This is not a defect to be prompted around; it is a property to design against, and the design response is to keep the list short rather than to make the items louder.

That means editing is the main activity, and the test for keeping a line is whether it is currently doing work. Most instruction documents contain three kinds of content: rules that matter now, rules added in response to a single incident that has not recurred, and general advice that was never actionable. The second and third categories crowd out the first, and removing them measurably improves adherence to what remains.

Specificity matters more than emphasis. Always verify before proceeding is not checkable and will be interpreted as a mood. Run the test suite and paste the failing output before making a second change is checkable, and the difference in compliance between the two is large. The reliable test is whether you could tell from the output that the instruction was followed. If you could not, the agent cannot either.

Placement matters and is easy to get wrong. Instructions in the middle of a long document compete with everything around them; instructions attached to the task they govern are followed better than the same words filed under general principles. Where something is critical, it belongs in the immediate brief rather than only in the standing document, and the small duplication is worth the adherence.

Prohibitions need a reason attached. Do not modify the schema is followed until a situation arises where modifying the schema is the obvious route to the goal, and then it becomes an obstacle to work around. Do not modify the schema because three other systems read it and will break survives that situation, because the reason is present when the conflict arises. This is the same reason recorded decisions need their rationale.

Finally, treat repeated non-compliance as information about the instruction. If the same rule is broken across different runs, the rule is ambiguous, conflicts with another rule, or is asking for something the situation does not permit. Restating it more firmly is the natural response and almost never works. Reading it as an ambiguity report and rewriting it is slower and does.

Every instruction you add is spent from the same budget, and the twentieth one is paid for by the nineteen that came before it.

Siddharth Sharma, Context Theory

Related questions

Do capital letters and words like 'must' help?

Marginally and unreliably, and the effect fades as the number of emphasised items grows, which it always does. Emphasis is a relative signal: if several instructions are emphatic, none is. It is worth reserving for one or two genuine invariants, at which point it is doing real work precisely because it is rare.

Should instructions be positive or negative?

Positive where a correct action exists, because a prohibition leaves the alternative unspecified and the alternative is where the surprise happens. Write the results to the review folder is better than do not write to the live folder, since the second is satisfied by any of a large number of places, only one of which you wanted.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide

2026 speed-to-lead benchmark · verified

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowInstruction adherence degrades as the instruction count rises and the items that get dropped vary between runs, so shortening the list improves compliance with what remains while adding emphasis to a long document does not.Testing compliance with one specific instruction placed in a short document and in the same document extended with additional rules.
ResponseAn instruction is followable in proportion to whether its execution would be visible in the output, so a rule that cannot be confirmed from the result is interpreted as a disposition rather than as a requirement.Asking, for each rule in a set, what in the output would demonstrate that it was followed.
SoftwareA prohibition without a reason is worked around once the prohibited route becomes the obvious path to the goal, whereas a prohibition carrying its consequence survives the conflict because the reason is present at the moment the conflict arises.Comparing adherence to a bare prohibition and to the same prohibition with its consequence stated, on a task where the prohibited route is the shortest one.
ConstraintRepeated non-compliance across separate runs indicates ambiguity, a conflict with another rule, or a requirement the situation cannot satisfy, none of which is addressed by restating the rule more firmly.Reading a repeatedly broken rule against the other rules in the same document for direct conflict.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one