Context Theory Get your growth audit

Answer

How many things should an agent do at once?

One, in the sense that its output makes a single claim. A run producing two results cannot be accepted or rejected.

One thing, meaning one reviewable result. A run that produces two unrelated outcomes forces whoever reviews it to accept both or reject both, and neither is the decision they wanted to make.

The unit is not the number of steps or tools but the number of claims the output makes. A run that reconciled the records and also tidied the naming has produced two results with one review attached, and the reviewer must accept the tidying to get the reconciliation or reject the reconciliation to refuse the tidying. This is the practical cost of unfocused runs and it appears at review rather than during execution.

The same argument applies to the multi-part task that arrives as one request. Handle the enquiries and update the report and check the outstanding items is three jobs, and running them together produces a report mixing three sets of findings, three sets of failures and a coverage figure that means nothing. Splitting them costs three briefs and produces three things that can each be judged.

There is a genuine efficiency argument for combining, and it is smaller than it appears. Work that shares context — the same records, the same investigation, the same setup — is cheaper done together, and that saving is real. What it buys is tokens and time; what it costs is reviewability, and reviewability is usually the constraint. Combining is right where the outputs will be accepted or rejected together anyway, which is a narrower set than it sounds.

Within a single run, doing several things simultaneously is a different question with a simpler answer. Parallel tool calls that gather independent information are straightforwardly good — reading four files at once rather than in sequence costs nothing extra and saves time. What should not be parallel is anything where one result should inform the next, which is a dependency mistaken for an opportunity.

The signal that a run did too much is available before any reading: the size of the result. A change touching eleven files for a task that should have touched two, a report covering four subjects, a summary with three unrelated sections. Each of these is visible immediately and is more reliable than any assessment of whether the work was appropriate, because it does not require knowing what the work was.

One qualifier. A run that discovers something outside its task should report it, not ignore it, and reporting is not doing. The discipline is that the additional finding arrives as information while the run's output remains a single claim, which preserves both the reviewability and the value of having noticed.

A run that did two things produces one decision, and the decision is wrong for at least one of them.

Siddharth Sharma, Context Theory

Related questions

Is it wasteful to run three times over the same material?

Somewhat, and the waste is tokens while the saving is review time, which is usually the more expensive of the two. Where the material is genuinely large and the setup genuinely costly, one run producing three separately labelled and separately acceptable outputs is a reasonable compromise, provided each can actually be accepted alone.

What about a task that naturally has several outputs?

If they stand or fall together — an analysis and the chart of it, a change and its test — that is one claim in two artefacts and belongs in one run. The test is whether a reviewer could sensibly accept one and reject the other. If they could, it is two things.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide

2026 speed-to-lead benchmark · verified

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowThe unit is the number of claims the output makes rather than the number of steps taken, because a run producing two unrelated results attaches one review decision to both and forces acceptance or rejection of each with the other.Identifying whether a reviewer could sensibly accept part of a run's output and reject the rest.
Buying behaviourCombining work that shares context saves tokens and time and costs reviewability, and reviewability is usually the binding constraint, which narrows the cases where combining is correct to those whose outputs would be accepted together anyway.Comparing the token saving from combining against the review time for the combined result.
SoftwareParallel tool calls gathering independent information cost nothing extra and save time, while parallelising steps where one result should inform another is a dependency mistaken for an opportunity.Checking whether any parallelised call requires the output of another to be correct.
ResponseResult size relative to task scope signals overreach before any reading and does not require knowing what the work was, which makes it more reliable than an assessment of appropriateness.Comparing the number of files, records or subjects in a result against the number the task required.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one