Context Theory Get your growth audit

Answer

How do you measure whether an automation is working?

Count things that used to be lost. Time saved is the metric suppliers propose and the one you cannot defend.

By counting work that used to be lost, against a baseline captured before it went live. Enquiries answered, calls picked up, quotes sent inside a target. Hours saved cannot be defended, because nobody logs the week that did not happen.

Automation measurement fails in two specific ways and both are avoidable at the start. The first is that no baseline was captured, so there is nothing to compare against and every claim about improvement is an assertion. The second is that the chosen metric is one nobody can check, which means the automation's fate depends on whether people happen to like it.

The baseline problem is the more damaging and the easier to fix. Before anything is switched on, record for a couple of weeks: how many enquiries arrived on each channel, how many got a response, how long that took, and how many got nothing. It is unglamorous data collection and it is the only thing that later distinguishes an automation that worked from one people are used to. Businesses that skip it are choosing, in advance, to be unable to answer the question.

On metrics, the useful ones share a property: they count events that either happened or did not. Enquiries that received a response within a target. Calls answered by something rather than ringing out. Quotes sent inside a stated window. Appointments confirmed without a person intervening. Each is checkable by someone who was not involved and none depends on anybody's impression.

Time saved fails that test, which is why it should not be the headline even though it is what suppliers propose. The hours are real and they are diffuse — a few minutes here, an interruption avoided there — and nobody logs the version of the week that did not happen. An estimate assembled after the fact is unfalsifiable in both directions, so it convinces people who already believed and nobody else. Use it as supporting colour if you like, never as the case.

There is a category of measurement that matters more than either and is almost always omitted: the harm check. An automation can improve its own metric while making something worse elsewhere, and the metric will not show it. A faster acknowledgement rate alongside a rising complaint rate. More appointments booked alongside more no-shows because the booking path stopped qualifying. More conversations started alongside a falling conversion rate. Pick one or two counter-metrics before go-live and watch them with the same seriousness as the primary one.

Finally, the measurement window has to match the business's own cycle rather than a billing period. Where an enquiry takes a quarter to become revenue, judging at thirty days measures activity and calls it success. That is the mechanism by which businesses conclude an automation works when it has produced more conversations and no more customers — and it is the same error that makes short trials favour whichever change generates the most immediate motion.

An automation with no baseline behind it cannot be proved to work, which means it cannot be defended, which means it will be removed by whoever inherits it.

Answer Production Engine, Context Theory

Related questions

What if we already switched it on without a baseline?

Reconstruct what you can and be explicit about the limits. Channel-level records often survive even when nobody was reading them — call logs, form submission timestamps, message histories — and a partial baseline stated honestly is worth more than a confident claim with nothing behind it. What you cannot reconstruct is what got no response, which is usually the number that mattered.

Should we run a proper A/B test?

Rarely worth it at small-business volume, and often actively misleading, because the sample needed to detect a realistic effect exceeds what most businesses see in a reasonable period. A before-and-after comparison on a stable count, with a counter-metric watched alongside, is weaker evidence in theory and considerably more usable in practice. Where a split is feasible, split by channel rather than by customer, since it avoids treating two people in the same household differently.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Firms that never responded to a web enquiry at all23%Category-wide
Teams responding to an inbound lead within 5 minutes7%Category-wide
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide

Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified

2026 speed-to-lead benchmark · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowA defensible automation metric counts events that either occurred or did not — responses within a target, calls answered rather than rung out, quotes sent inside a window — and is checkable by someone who was not involved in the project.Whether each proposed metric can be recomputed from system records by a person with no knowledge of the automation.
WorkflowTime saved is unfalsifiable because the hours are diffuse and nobody records the version of the period that did not happen, so an after-the-fact estimate persuades only those already convinced.Any attempt to reconstruct the claimed hours from logged activity records rather than from recollection.
WorkflowAn automation can improve its own metric while degrading something adjacent — faster acknowledgement with rising complaints, more bookings with more non-attendance — so counter-metrics must be selected before go-live rather than investigated after.The counter-metric values recorded in the same baseline period as the primary measure.
ProcurementJudging an automation over a period shorter than the business's own enquiry-to-revenue cycle measures activity rather than outcome, which is how more conversations without more customers is recorded as success.The business's own median elapsed time from first enquiry to closed revenue, compared with the evaluation window used.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one