Context Theory Get your growth audit

Answer

How do you know when an AI workflow has degraded?

By comparing against a fixed set of cases and by watching the distribution of what arrives. Nothing about degradation announces itself.

By re-running a fixed set of cases and comparing, and by watching the distribution of inputs and outputs for shifts. Degradation produces no error, so nothing surfaces it unless something is measuring.

Degradation has several causes and one property: it is silent. The inputs change shape, an upstream system starts sending a new format, a model version moves, a document the workflow depends on is updated, the business starts doing something slightly different. In every case the workflow keeps running, keeps producing well-formed output, and keeps reporting success. Nothing in the ordinary operation of the system distinguishes the before from the after.

The primary detector is a fixed evaluation set re-run on a schedule. Because the cases and the correct answers do not move, any change in the result is a change in the system or its dependencies. This is the only signal that isolates the workflow from everything around it, and it is why the set has to stay fixed: an evaluation set that grows between runs cannot tell you whether the difference was the system or the set.

The second detector is the input distribution, and it usually moves first. What proportion of items fall into each category, how long they are, which fields are populated, where they come from. A shift here means the workflow is now handling something other than what it was built for, and it will keep handling it confidently. Watching a handful of these counts costs almost nothing and gives the earliest warning available.

The third is the output distribution, which is the mirror. A classifier whose share of items in one category jumps, a workflow whose refusal rate falls to zero, a process whose average output length changes: each indicates something moved even when the aggregate accuracy has not been remeasured. These are cheap to compute continuously and they catch the cases where the input looks the same and the behaviour is not.

The fourth is cost per item, which is the most underused. Retries, longer contexts and additional tool calls all show up here before they show up in quality, and a rising unit cost with stable output is a reliable indication that the system is working harder for the same result. This one requires no labels and no evaluation set, which makes it the easiest to put in place.

The fifth is downstream correction. Wherever a person fixes an output, records a complaint or rejects a record, that is a measurement being generated and usually discarded. Capturing it converts everyone who already notices errors into a detection layer, and the resulting series is the closest thing to a continuous quality measure that most businesses can obtain.

A degraded workflow looks exactly like a working one from the outside, which is why the question has to be asked on a schedule rather than when something seems wrong.

Siddharth Sharma, Context Theory

Related questions

What is the most common cause of degradation?

A change in what arrives rather than a change in the system. Businesses evolve, upstream formats shift, and a new category of item starts appearing that the workflow was never designed for and will classify anyway. This is why the input distribution is worth watching even when nothing about the workflow has been touched.

Should you be notified when a model version changes?

Where the platform permits it, yes, and it should trigger a re-run of the evaluation set rather than a decision. A version change is not evidence of a problem and it is a reason to measure, and coupling the two means the question gets asked at the moment it becomes relevant rather than at the next scheduled interval.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Firms that never responded to a web enquiry at all23%Category-wide

2026 speed-to-lead benchmark · verified

Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowDegradation from shifted inputs, changed upstream formats, updated dependencies or a moved model version all share the property of producing well-formed output and a success report, so ordinary operation does not distinguish before from after.Changing an input characteristic the workflow does not handle and observing whether any part of the system reports anything.
SoftwareA fixed evaluation set is the only detector that isolates the workflow from its surroundings, because unchanged cases and answers mean any change in result is attributable to the system or its dependencies.Comparing evaluation results across runs where the set was held fixed against runs where cases were added.
ResponseThe input distribution moves before quality does, so tracking category shares, item lengths, populated fields and sources provides the earliest available warning at negligible cost.Comparing the current month's input mix against the mix at the time the workflow was built.
Buying behaviourCost per item rises through retries, longer contexts and additional tool calls before quality falls, and it requires no labels or evaluation set, which makes it the cheapest degradation signal to instrument.Plotting period spend divided by items handled across several months.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one