Context Theory Get your growth audit

Answer

What happens when an AI agent has too much context?

It gets vaguer rather than slower. Specific instructions stop being followed and present facts stop being used, with no error anywhere.

It degrades quietly. Specific instructions get followed approximately, facts that are present stop being used, and early material loses influence to recent material. Nothing errors, and the output stays fluent throughout.

The expectation is that too much context produces an error or a refusal. What actually happens is worse for practical purposes: the work continues and gets less precise. The named symptoms are consistent enough to be diagnostic. Instructions that were followed exactly earlier are followed approximately. A fact that is demonstrably present in the material is not used. The agent answers from general knowledge where it should be answering from the supplied document. Its summaries become plausible rather than particular.

The mechanism is that attention is a finite resource spread across everything present. Every additional token competes with every other, so the relative weight of any specific instruction falls as the total rises. This is why emphasis does not rescue a bloated context: emphasis is relative, and it is competing against a larger field. It is also why the effect appears gradually rather than at a threshold, which is what makes it hard to attribute.

There is a second, distinct effect in long sessions: recency. Material introduced recently carries more influence than material introduced at the start, so the original brief becomes progressively less governing while the last few exchanges become more so. An agent forty steps into a job is substantially shaped by step thirty-nine and only lightly by the instruction that set the task, which is the mechanism behind drift.

The practical diagnosis is straightforward once you know to look. Ask a question whose answer is unambiguously in the supplied material. If the answer is wrong, hedged, or drawn from general knowledge, the working set is too large regardless of what any limit says. This test takes seconds and is more informative than any token count, because it measures the property you care about rather than the one that is easy to display.

The remedy is not compression, which loses detail nobody chose. It is restarting from a written state: take the durable outcome of the work so far, put it in a file, begin a new session with that file and the current task. Everything that mattered survives because it was written down, and everything that was competing for attention does not. Teams that adopt this describe the improvement as the model getting better, which is a reasonable description of what it feels like and a misleading account of what changed.

An overloaded agent does not fail; it becomes generic, which is the failure mode hardest to notice and easiest to accept.

Siddharth Sharma, Context Theory

Related questions

Does a larger context window remove the problem?

It moves the point at which it appears and does not change its nature, because the degradation is about how much competes for attention rather than about hitting a wall. A larger window makes it possible to hold more, which in practice means more gets held, and the symptom arrives later in the session at a similar level of dilution.

How do you tell dilution from a model limitation?

Run the same request in a fresh session with only the relevant material. If it succeeds there and fails in the long session, the difference is the context rather than the capability. This is worth doing before concluding a task is beyond the system, because the conclusion is expensive and wrong more often than people expect.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Visibility lift in AI-generated answers from GEO methodsup to 40%Category-wide
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide

Aggarwal et al., "GEO: Generative Engine Optimization", Princeton / Georgia Tech / IIT Delhi / Allen Institute for AI — KDD 2024 · GEO-bench · 10,000 queries across 8 domains · verified

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

What is specific to this page.

Evidence
Kind Claim Check it against
SoftwareExcess context degrades output quality rather than producing an error, with instructions followed approximately and present facts left unused, so no layer of the system reports the condition and the prose remains fluent throughout.Asking a question whose answer is unambiguously present in supplied material at increasing total context volumes.
WorkflowAttention is distributed across everything present, so the relative weight of any specific instruction falls as total volume rises, which is why emphasis does not compensate — emphasis is a relative signal competing in a larger field.Comparing adherence to an emphasised instruction in a short context and the same instruction in a long one.
ResponseRecency is a separate effect from volume: material introduced late carries more influence than the original brief, so an agent deep into a run is shaped more by its recent steps than by the instruction that set the task.Asking an agent late in a long run to restate its objective and comparing it with the brief it was given.
Buying behaviourRestarting from a written state outperforms compressing the session, because compression discards detail that nobody selected while a written state preserves the outcomes deliberately and leaves the competing material behind.Comparing continuation after compaction against continuation from a written state file on the same interrupted task.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one