Answer
When should an AI agent decide on its own?
When the decision is reversible, internal to the work, and made from information it can actually see.
When the choice can be undone, its effects stay inside the work, and everything needed to make it is visible to the agent. If any of those three fails, the decision belongs to a person however small it looks.
Autonomy is usually configured as a level — this agent is supervised, that one is not — and that framing does not survive contact with real work, because a single run contains decisions of wildly different consequence. The useful unit is the decision, and three properties determine whether it can be made without asking.
Reversibility is the first and the most important. A choice that can be undone in one command has a bounded cost even when it is wrong, and bounded cost is what makes delegation rational. This is why the practical route to more autonomy runs through version control, snapshots and soft deletion rather than through better judgement: making things reversible expands the set of decisions that can safely be delegated without changing anything about the agent.
Containment is the second. A decision whose effects stay inside the work — a structure, an ordering, an intermediate format, an approach — can be re-decided later at the cost of the work itself. A decision that leaves the work becomes a fact other things depend on: a record other systems read, a message someone received, a name that gets used elsewhere. The leaving is what makes it consequential, not the difficulty.
Information access is the third and the most frequently overlooked. An agent can only decide well from what it can see, and the decisions that go wrong are usually those requiring context that exists in somebody's head: which customer is sensitive, which record is authoritative, why the previous approach was abandoned. These look routine from the inside, which is exactly why they get made autonomously and exactly why they are wrong.
Applying the three gives a division most teams find surprisingly liberal. Nearly all execution decisions pass: how to structure the work, which order to do things in, what to call an intermediate file, how to phrase a draft, which approach to try first. Nearly all commitments fail: prices, dates, scope, anything sent, anything deleted, anything another party will rely on. The uncomfortable middle is small and it is where the design attention belongs.
One thing that should not appear in this calculation is how well the agent has performed recently. A run of good decisions is evidence about the decisions that came up, not about the ones that have not yet. Expanding autonomy is a reasonable thing to do as evidence accumulates, and the evidence that should move it is structural — a rollback now exists, a check now runs, the information is now written down — rather than a record of nothing having gone wrong.
Autonomy should be granted per decision rather than per agent, because the size of a decision is a poor guide to the size of its consequences.
Siddharth Sharma, Context Theory
Related questions
Should an agent decide when to ask?
It should be able to ask at any time and it should not be the thing that determines which decisions require asking. Leaving the boundary to the agent's judgement means the boundary is set by whether the situation felt significant, and the failures are precisely the cases that did not feel significant. Enforce the required stops structurally and let voluntary questions be a bonus.
What about decisions that are reversible but embarrassing?
Treat visibility as a form of irreversibility, because it is. A message someone has read cannot be unread, even if the record can be corrected, and the same applies to anything published, sent or shown to a third party. This is why audience is a separate category in most sensible permission policies rather than being folded into technical reversibility.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
| Firms that never responded to a web enquiry at all | 23% | Category-wide |
2026 speed-to-lead benchmark · verified
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Constraint | Autonomy configured per agent rather than per decision fails because a single run contains decisions of widely differing consequence, so the unit of the grant must be the decision and not the system making it. | Listing the distinct decisions in one representative run and classifying each by consequence. |
| Procurement | Making actions reversible expands the set of decisions that can be delegated without altering the agent at all, which is why version control, snapshots and soft deletion are autonomy mechanisms rather than only recovery mechanisms. | Counting which currently gated decisions would pass a reversibility test after a rollback path is added. |
| Workflow | Decisions requiring context that exists only in a person's knowledge appear routine from within the work, which is why they are the ones made autonomously and the ones most often wrong. | Reviewing past incorrect autonomous decisions for whether the missing information was recorded anywhere the agent could read. |
| Buying behaviour | A record of recent correct decisions is evidence about the decisions that arose rather than about those still to come, so autonomy should widen on structural changes such as an added rollback or check rather than on an absence of incidents. | Naming what changed in the system since the last boundary expansion, other than elapsed time without failure. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one