Context Theory Get your growth audit

Answer

How do you build an evaluation set for an AI workflow?

From real inputs you did not choose, with the correct answers established once and written down. Invented cases test your imagination.

Take a period of real inputs without selecting them, establish the correct answer for each once, and add every production failure afterwards. Invented cases test the situations you thought of, which are not the ones that fail.

The first decision is where the cases come from, and it settles most of the quality of the result. Real inputs, taken from a period without filtering, contain the distribution the workflow actually faces: the malformed ones, the unusual ones, the ones with a field missing, the ones that are two requests in one message. Cases written by the person who built the workflow contain the situations that occurred to them, which by construction excludes the ones that will cause trouble.

The second is establishing the correct answer, which is the expensive part and is a one-off. For each case, what should the workflow have produced. This has to be decided by someone who knows the business rather than by the system being evaluated, and the act of deciding frequently surfaces that the correct answer is contested — two people in the business disagree about what should happen. That disagreement is a finding worth more than the evaluation, because the workflow cannot be right about something the business has not decided.

The third is size, and the answer is smaller than people fear. A few dozen well-chosen real cases with reliable labels will tell you most of what you need, particularly at the start, and is achievable in an afternoon. A large set with unreliable labels is worse than a small careful one, because the errors in the labels become the ceiling on what the evaluation can tell you.

The fourth is growth, and this is what makes the set valuable over time. Every failure found in production becomes a case. This costs a minute at the moment of discovery and it means the same failure cannot silently return, which is exactly the guarantee a regression suite provides in ordinary software. Sets that grow this way come to represent the real failure surface within a few months, without anyone designing them.

One thing to guard against: the set becoming unrepresentative as it grows. Adding every failure and no successes gradually turns an evaluation set into a collection of hard cases, and the rate measured over it stops resembling the rate in production. Keeping the original sampled portion intact and separate, and reporting both numbers, preserves the meaning of each.

Finally, keep the set where the workflow lives and under the same version control. An evaluation set on someone's machine is a set that stops being run, and a set that is not run is a document about a measurement that no longer takes place.

A test set you wrote is a test of your imagination; a test set you sampled is a test of the system.

Siddharth Sharma, Context Theory

Related questions

Can a model generate evaluation cases?

It can generate variations of cases you already have, which is useful for testing robustness to phrasing, and it cannot generate the distribution of real inputs because it does not know it. Generated cases are a supplement to sampled ones and a poor substitute, since the point of sampling is to capture what nobody would think to write.

What about sensitive data in evaluation cases?

Use the real structure with the identifying content replaced, and be careful that the replacement does not remove what made the case hard. A record anonymised into tidiness is no longer the awkward case it was chosen for. Where anonymisation would destroy the difficulty, the case belongs in a set held under the same controls as the data itself.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide

2026 speed-to-lead benchmark · verified

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowUnfiltered real inputs contain malformed, unusual and compound cases that a set written by the workflow's author excludes by construction, because authored cases contain only the situations that occurred to the author.Comparing the input characteristics of a sampled period against a hand-written test set for the same workflow.
ResponseEstablishing correct answers frequently reveals that people in the business disagree about what should happen, which is a more valuable finding than the evaluation because a workflow cannot be correct about an undecided question.Having two people independently label the same set of cases and comparing their answers.
SoftwareLabel reliability rather than set size bounds what an evaluation can establish, so a small set with dependable labels outperforms a large one whose label errors set the ceiling on measurable accuracy.Re-labelling a sample of an existing evaluation set and measuring the disagreement rate.
ConstraintAdding production failures without adding successes turns an evaluation set into a collection of hard cases, so the measured rate ceases to resemble production and the sampled portion must be kept separate and reported alongside.Comparing the measured rate over the full grown set against the rate over the original sampled portion.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one