Context Theory Get your growth audit

Answer

How do you stop an AI agent repeating the same mistake?

Change something outside the agent. A correction inside a session dies with it, and a new rule rarely lands.

Change the setup, not the session. A correction given in conversation disappears when the session does. The durable fixes are a check that catches the mistake, better material, or removing the tool that made it possible.

The recurrence has an obvious cause that is easy to miss while you are in the middle of correcting it: the correction lives in a session, and the session ends. The next run starts from the same brief, the same material and the same tools that produced the mistake, so it produces the mistake. Nothing about this is surprising once stated, and it explains why teams find themselves giving the same correction weekly.

There are four places a durable fix can live, and they are in ascending order of reliability. In the instructions, as a rule. In the material, as a fact that removes the ambiguity. In a check, which catches the mistake whether or not it was avoided. Or in the tools, where the option that permits the mistake no longer exists. Most teams reach only for the first, which is the weakest, because it is the one that feels like an answer.

A rule works when the mistake was a genuine matter of preference — a convention, a format, an ordering. It works poorly when the mistake came from a false belief about the situation, because the rule then competes with the belief and the belief is generated fresh each run. If the agent keeps assuming a field is populated, the fix is not a rule saying it may be empty; it is material showing that it frequently is, or a check that fails when the assumption is made.

A check is the most reliably underused of the four. A validation, a test, an assertion, a query that must return nothing: any of these converts a recurring mistake from something that must be avoided into something that cannot pass. This is strictly better than prevention, because it works on the runs where the instruction was not attended to, which are the runs where the mistake happens.

Tool removal is the strongest and is available more often than people assume. If an agent keeps modifying files it should not, the fix may be that it cannot reach them. If it keeps using a general command where a specific one is required, removing the general one ends the question. This is uncomfortable because it feels like a blunt instrument, and it is a blunt instrument that works.

One diagnostic worth running first. Before fixing anything, check whether the mistake is actually the same mistake. Three superficially similar failures often have different causes — one from ambiguous material, one from a missing tool result, one from a genuinely wrong instruction — and a single fix applied to all three resolves one and leaves two, which then looks like the fix not working.

Telling an agent it was wrong changes this run and nothing else, which is why the same mistake arrives again on Tuesday looking brand new.

Siddharth Sharma, Context Theory

Related questions

Does adding it to the project instructions count as a durable fix?

It is durable in the sense that it persists, and it is weak in the sense that it competes with everything else in that document. As instruction files grow, the marginal rule is followed less reliably than the ones already there, so a rule added to a long file is a fix that decays. If the instruction file is where the fix must go, something usually needs to come out.

What if the mistake only happens sometimes?

Intermittent recurrence usually means the trigger is in the input rather than in the agent. Find what is different about the runs where it happens — an unusual record, an empty field, a longer document, a different route through the work — and the fix is nearly always to handle that case explicitly rather than to make the general behaviour more careful.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Agents who give up after one contact44%Category-wide

2026 speed-to-lead benchmark · verified

Multi-study aggregate · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowA correction delivered in conversation does not persist beyond the session, so the next run begins from the same brief, material and tools that produced the mistake, which is why the same correction recurs on a weekly cadence in practice.Recording where each correction was made and whether it survived into a fresh session.
SoftwareA rule addresses a mistake arising from preference but not one arising from a false belief about the situation, because the belief is regenerated each run and the rule competes with it rather than replacing it.Classifying a recurring mistake by whether it reflects a convention or an assumption about the data, and testing a rule against each.
ResponseA check is superior to prevention because it operates on the runs where the instruction was not attended to, which are precisely the runs in which the mistake occurs, whereas a rule is only effective on runs that would have been compliant anyway.Comparing recurrence rates under an added rule against under an added validation for the same mistake.
ConstraintSuperficially identical failures frequently have distinct causes — ambiguous material, an unread tool result, a wrong instruction — so a single fix resolves one and leaves the others, which presents as the fix having failed.Tracing each instance of a recurring failure to its immediate cause before selecting a fix.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one