Context Theory Get your growth audit

Answer

How do you know whether an autonomous agent is actually working?

Measure the outcome the agent exists to produce, not the activity it generates. Activity is the easiest thing to sustain while failing.

Measure the outcome it was built to change, the share of cases it covered, how often its output needed correcting, and what it cost. Run counts and output volume can all look healthy while the work is going wrong.

The dashboard that gets built first shows runs completed, items processed, and no errors. All three can be perfectly healthy while the system is producing wrong output at scale, because none of them looks at the content. This is the specific hazard of autonomous work: the failure signal that a human process would generate — someone noticing, someone asking, someone getting stuck — has been removed, and nothing was installed in its place.

Start from the outcome the agent was supposed to change. If it triages enquiries, the outcome is how quickly enquiries get a substantive reply and how many fall through. If it reconciles records, the outcome is how many discrepancies reach month end. If it drafts, the outcome is how much editing the drafts need. That number should have moved, and if it has not, the system is running rather than working. This is the measurement nobody wants to define because it can fail.

The second measure is coverage: of everything the agent was meant to handle, what proportion did it actually handle, and what happened to the rest. Autonomous systems fail silently at the edges long before they fail in the middle, and a quietly growing pile of unhandled cases is the most common form of degradation. Coverage is also the measure that catches an agent narrowing its own scope, which happens gradually and never gets reported.

The third is correction rate, which requires that corrections be recorded rather than made. In most businesses someone quietly fixes the output and moves on, and the resulting impression that things are fine is unfalsifiable. Even a crude record — a tick when something needed changing — converts an impression into a series, and a series will show a trend that no amount of recollection ever does.

The fourth is cost per unit of outcome, not cost in total. Total spend rises with volume and tells you nothing; cost per handled case tells you whether the system is becoming less efficient, which is usually the earliest quantitative sign that inputs have shifted or that the agent is retrying more than it did.

Finally, sample the actual output on a schedule. Not to catch individual errors, which is not worth the time, but because the numbers above are all proxies and every proxy eventually detaches from the thing it stands for. A handful of real cases read properly each week is the only measurement that sees what the customer sees, and it is the one that gets dropped first when everything appears to be going well.

A failing agent produces the same amount of activity as a working one, which is why activity is the one thing not worth watching.

Siddharth Sharma, Context Theory

Related questions

What is the earliest sign that something has gone wrong?

Usually a change in the shape of the inputs rather than in the behaviour of the agent. The system keeps doing what it did while what arrives has changed, and the mismatch shows up as a rising share of cases routed to a catch-all category or handled with unusual speed. Watching the distribution of outcomes is more informative than watching any single rate.

How do you measure an agent whose output is judgement?

By what happens downstream, since the output itself has no external referent. Whether the recommendation was acted on, whether the draft was sent largely intact, whether the analysis changed a decision. These are weaker measurements than a reconciliation and they are real, and they are considerably better than asking whether people feel it is helping.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Firms that never responded to a web enquiry at all23%Category-wide
Average B2B first-response time42Category-wide

Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · hours · 1.25M inbound leads across 2,241 US firms · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowRun counts, item counts and error rates can all be healthy while output is wrong, because none of them inspects content, and autonomous operation has already removed the human friction that would otherwise surface the problem.Checking whether the existing monitoring for an autonomous process would change if its outputs became systematically incorrect.
ResponseAutonomous systems degrade at the edges before the middle, so coverage — the share of intended cases actually handled and the disposition of the remainder — detects a narrowing scope that is never self-reported.Comparing the count of cases the agent handled against the count of cases that entered its intended scope.
ConstraintCorrections made without being recorded produce an unfalsifiable impression of adequacy, so even a binary record of whether an output needed changing converts recollection into a series in which a trend becomes visible.Introducing a single correction flag on reviewed output and comparing the first month's rate against the prevailing impression.
Buying behaviourCost per handled case rather than total spend is the diagnostic quantity, because total cost tracks volume while unit cost rises when inputs shift or retries increase, which is usually the earliest quantitative sign of degradation.Dividing period spend by cases handled and comparing the series across months.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one