Context Theory Get your growth audit

Answer

How much should you trust an AI system that has been right so far?

As much as the inputs it has seen justify. A clean record is evidence about those cases, not about the next one.

As far as the inputs it has actually seen justify, and no further. A run without errors is evidence about those cases; the ones that will cause trouble are by definition the ones that have not arrived yet.

The reasoning that feels natural here is inductive: it has been right two hundred times, so it will probably be right the next time. That holds where the next case is drawn from the same distribution, and the cases that cause problems are precisely the ones that are not — the unusual customer, the malformed input, the situation nobody anticipated. A record of success on ordinary inputs is only weak evidence about extraordinary ones.

There is a second problem with the record itself, which is that it is usually not a record. Most businesses do not track corrections; someone fixes an output and moves on, and the impression of correctness is built from the absence of memorable failures rather than from data. That impression is systematically optimistic, because the small corrections that were made are exactly the ones nobody remembers.

The practical translation is that trust should follow evidence rather than time. Widening what a system does unsupervised is reasonable when something structural has changed — a check now runs, a rollback path exists, the failure modes have been measured, the information the system was missing is now written down. It is not reasonable merely because nothing has gone wrong, since nothing going wrong is also what a system that has not yet met a hard case looks like.

There is a specific hazard as reliability improves, and it applies to people rather than to systems. A checking habit that usually finds nothing erodes, because the effort has no visible return. The errors that remain are then more likely to reach a customer, precisely because nothing was looking. This is why the response to an improving system should be to check faster rather than to check less, and why removing a check is a decision that deserves the same scrutiny as adding an autonomous capability.

The useful question to ask about a well-performing system is not how much to trust it but what has it not seen. Seasonal peaks, an unusual category of request, an edge case in the data, a change in an upstream format. Those are enumerable in an afternoon for most workflows, and knowing which have not yet occurred tells you far more about the risk ahead than any number of successful runs behind.

None of this is an argument for permanent suspicion. A system whose failure modes have been measured, whose errors are caught by something, and whose consequences are bounded deserves to be given more to do. The argument is only about what constitutes the evidence for that, and a clean history is the weakest form of it.

A clean record tells you what the system did with the easy months, which is the information you needed least.

Siddharth Sharma, Context Theory

Related questions

How long before a track record means something?

It is about coverage rather than duration. A month that included a quarter-end, an unusual request type and a system outage tells you more than a year of ordinary Tuesdays. The question to answer is which of the situations you care about have actually occurred during the observation period.

What should trigger a review of a system that is working well?

A change in inputs, a platform change, a change in what the output is used for, and the passage of a fixed interval regardless. The last one matters because the first three are not always noticed, and a scheduled review is the only mechanism that does not depend on someone spotting a change that nobody announced.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide

2026 speed-to-lead benchmark · verified

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowA record of correct outputs is evidence about the input distribution encountered, and the cases that cause problems are those outside it, so success on ordinary inputs is weak evidence about extraordinary ones.Listing the input situations the workflow has not yet encountered during its observed period.
ConstraintThe perceived track record is usually an impression rather than a record, because corrections are made without being logged, and the small fixes that were performed are the ones least likely to be remembered.Comparing the recorded correction rate against the prevailing impression of accuracy.
ResponseA checking habit that consistently finds nothing erodes because the effort has no visible return, so the residual errors become more likely to reach a customer exactly as the system improves.Comparing the proportion of outputs actually checked in the first month of operation against the current month.
Buying behaviourEnumerating which situations a workflow has not yet met — seasonal peaks, unusual categories, data edge cases, upstream format changes — is achievable quickly and is more informative about forward risk than the count of successful runs.Listing the workflow's known input situations and marking which have occurred since deployment.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one