Answer
How do you measure whether an automation is working?
Count things that used to be lost. Time saved is the metric suppliers propose and the one you cannot defend.
By counting work that used to be lost, against a baseline captured before it went live. Enquiries answered, calls picked up, quotes sent inside a target. Hours saved cannot be defended, because nobody logs the week that did not happen.
Automation measurement fails in two specific ways and both are avoidable at the start. The first is that no baseline was captured, so there is nothing to compare against and every claim about improvement is an assertion. The second is that the chosen metric is one nobody can check, which means the automation's fate depends on whether people happen to like it.
The baseline problem is the more damaging and the easier to fix. Before anything is switched on, record for a couple of weeks: how many enquiries arrived on each channel, how many got a response, how long that took, and how many got nothing. It is unglamorous data collection and it is the only thing that later distinguishes an automation that worked from one people are used to. Businesses that skip it are choosing, in advance, to be unable to answer the question.
On metrics, the useful ones share a property: they count events that either happened or did not. Enquiries that received a response within a target. Calls answered by something rather than ringing out. Quotes sent inside a stated window. Appointments confirmed without a person intervening. Each is checkable by someone who was not involved and none depends on anybody's impression.
Time saved fails that test, which is why it should not be the headline even though it is what suppliers propose. The hours are real and they are diffuse — a few minutes here, an interruption avoided there — and nobody logs the version of the week that did not happen. An estimate assembled after the fact is unfalsifiable in both directions, so it convinces people who already believed and nobody else. Use it as supporting colour if you like, never as the case.
There is a category of measurement that matters more than either and is almost always omitted: the harm check. An automation can improve its own metric while making something worse elsewhere, and the metric will not show it. A faster acknowledgement rate alongside a rising complaint rate. More appointments booked alongside more no-shows because the booking path stopped qualifying. More conversations started alongside a falling conversion rate. Pick one or two counter-metrics before go-live and watch them with the same seriousness as the primary one.
Finally, the measurement window has to match the business's own cycle rather than a billing period. Where an enquiry takes a quarter to become revenue, judging at thirty days measures activity and calls it success. That is the mechanism by which businesses conclude an automation works when it has produced more conversations and no more customers — and it is the same error that makes short trials favour whichever change generates the most immediate motion.
An automation with no baseline behind it cannot be proved to work, which means it cannot be defended, which means it will be removed by whoever inherits it.
Answer Production Engine, Context Theory
Related questions
What if we already switched it on without a baseline?
Reconstruct what you can and be explicit about the limits. Channel-level records often survive even when nobody was reading them — call logs, form submission timestamps, message histories — and a partial baseline stated honestly is worth more than a confident claim with nothing behind it. What you cannot reconstruct is what got no response, which is usually the number that mattered.
Should we run a proper A/B test?
Rarely worth it at small-business volume, and often actively misleading, because the sample needed to detect a realistic effect exceeds what most businesses see in a reasonable period. A before-and-after comparison on a stable count, with a counter-metric watched alongside, is weaker evidence in theory and considerably more usable in practice. Where a split is feasible, split by channel rather than by customer, since it avoids treating two people in the same household differently.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Firms that never responded to a web enquiry at all | 23% | Category-wide |
| Teams responding to an inbound lead within 5 minutes | 7% | Category-wide |
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
2026 speed-to-lead benchmark · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | A defensible automation metric counts events that either occurred or did not — responses within a target, calls answered rather than rung out, quotes sent inside a window — and is checkable by someone who was not involved in the project. | Whether each proposed metric can be recomputed from system records by a person with no knowledge of the automation. |
| Workflow | Time saved is unfalsifiable because the hours are diffuse and nobody records the version of the period that did not happen, so an after-the-fact estimate persuades only those already convinced. | Any attempt to reconstruct the claimed hours from logged activity records rather than from recollection. |
| Workflow | An automation can improve its own metric while degrading something adjacent — faster acknowledgement with rising complaints, more bookings with more non-attendance — so counter-metrics must be selected before go-live rather than investigated after. | The counter-metric values recorded in the same baseline period as the primary measure. |
| Procurement | Judging an automation over a period shorter than the business's own enquiry-to-revenue cycle measures activity rather than outcome, which is how more conversations without more customers is recorded as success. | The business's own median elapsed time from first enquiry to closed revenue, compared with the evaluation window used. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one