Answer
How do you evaluate an AI workflow?
Run it against inputs whose answers you already know, and count. Everything else is impression management.
Run it over a set of real inputs whose correct outcomes you already know, and count how often it agrees. Demonstrations and sample outputs measure how the system performs when someone is watching and choosing the examples.
The question is usually answered with a demonstration: run it on a few cases, look at the output, form a view. That produces an impression that is systematically optimistic, because the cases were chosen by someone who knows the system and the assessment is made by someone who wants it to work. It also produces no number, so there is nothing to compare against next month.
The alternative is straightforward and rarely done. Take a period of real work whose outcomes are recorded — last month's enquiries, last quarter's invoices, the tickets from a fortnight — run the workflow across it, and compare its output to what actually happened. This yields a rate, a set of specific disagreements, and the ability to repeat the exercise after any change.
The disagreements are worth more than the rate. A workflow that is accurate overall but wrong on a recognisable class of input is a workflow with a routing rule waiting to be written, and that finding is only available from a real distribution. Cases assembled by hand contain the situations someone thought of; a real month contains the ones nobody did, which is where the failures live.
What to measure depends on the output, and the useful discipline is to measure the thing the workflow exists to change. A classifier is measured by agreement with the correct category, and separately by what happens to the cases it gets wrong. A drafting workflow is measured by how much editing the drafts need, which is observable if edits are recorded. A retrieval workflow is measured by whether each claim in the output is supported, not by whether relevant documents were found.
Two measures should always be present alongside accuracy. Coverage: what proportion of the intended cases the workflow handled at all, since silent skipping does not appear in an accuracy figure computed over handled items. And cost per item, which is what tells you whether the thing is drifting before accuracy does.
Finally, an evaluation is only worth building if it will be run again. The value comes from comparison — before and after a change, this month against last — and a single measurement taken once is a fact about a moment. Keeping the input set fixed, so results are comparable, is the part that makes the exercise cumulative rather than repeated.
An evaluation you can fail is a measurement; one you cannot is a demonstration with a spreadsheet attached.
Siddharth Sharma, Context Theory
Related questions
What if there is no recorded correct answer?
Then the first task is establishing one for a small set, which is real work and is unavoidable. Grading fifty cases by hand once gives you a reference that every subsequent evaluation runs against, and it is far cheaper than the alternative of never knowing. Businesses that skip this generally have opinions about their workflows and no evidence.
How often should an evaluation run?
After every change to the workflow, and on a schedule regardless of changes, because the inputs move even when the system does not. The schedule matters more than the frequency: a quarterly run that actually happens beats a weekly one that lapses after a month.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
| Odds of qualifying a lead — replying within the first hour vs after it | 7× | Category-wide |
2026 speed-to-lead benchmark · verified
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Demonstration-based assessment is systematically optimistic because the cases are selected by someone familiar with the system and judged by someone with an interest in the result, and it produces no figure to compare against later. | Comparing the outcome of a demonstration against a run over an unselected historical period. |
| Response | A real period of inputs contains the situations nobody anticipated while a hand-assembled set contains only those someone thought of, which is why the specific disagreements from a real distribution are worth more than the aggregate rate. | Comparing the failure types found in a historical run against those found in a constructed test set. |
| Software | Accuracy computed over handled items conceals silent skipping, so coverage must be measured separately, and cost per item moves before accuracy does, which makes it the earlier warning. | Comparing the count of items the workflow handled against the count that entered its intended scope. |
| Constraint | An evaluation's value comes from comparison over time, so keeping the input set fixed is what makes results cumulative, and a single measurement is a fact about one moment rather than a property of the workflow. | Attempting to compare two evaluation results computed over different input sets. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one