Context Theory Get your growth audit

Answer

How many tools should an AI agent have?

As few as the job needs. Every additional tool competes for the selection decision and widens what a misunderstanding can reach.

As few as the job actually needs. Each additional tool competes in the selection decision and widens the blast radius of a misunderstanding, and most agents perform better with a short, well-described set than a broad one.

Tool count is treated as a capability question and is really a reliability question. Selecting the right tool is a decision made from a list of descriptions, and like any decision made from a list, it degrades as the list grows and as the entries become harder to distinguish. An agent with a small set of clearly separated tools chooses correctly far more often than one with a large set containing several plausible candidates for the same job.

The second effect is on risk, and it is simpler. A capability that exists can be used in a situation nobody anticipated. Removing a tool the job does not require removes an entire class of unexpected behaviour at no cost, and it is the cheapest safety improvement available in most agent configurations because it requires no policy, no prompt and no review.

This is why connecting every available integration is a mistake even when each one individually seems harmless. The effects compound: more descriptions to distinguish between, more overlapping capabilities, more surface. Businesses that connect three things deliberately generally get better behaviour than those that connect fifteen because the connectors were offered, and the difference is visible in tool selection long before it shows up as an incident.

Given a set, the quality of the descriptions matters more than the count. A tool whose description says what it does, what it does not do, when to prefer it over the near neighbour, and what its output looks like is chosen correctly. A tool described in three words is chosen approximately. This is ordinary interface design, and the interface here has a machine on the other side of it: the same clarity that helps a new colleague pick the right one is what helps here, for recognisably similar reasons.

Overlap deserves specific attention because it is the most common defect in a grown tool set. Two tools that could both plausibly serve a request produce inconsistent behaviour across runs, and the inconsistency looks like flakiness rather than like a configuration problem. When a set has grown organically, the useful exercise is to find pairs where a reasonable reader could not say which one applies, and to merge or rename them.

There is a case for a larger set, and it is worth stating so this does not read as an argument for minimalism. Where a job genuinely spans several systems, giving the agent a way to reach each of them beats forcing it to work around a gap, and an agent that cannot reach something it needs will improvise, which is worse than a slightly crowded tool list. The rule is that each tool should be there because a step of the job needs it, not because it was available.

A tool an agent has and does not need is not neutral; it is a wrong turning that has been made available.

Siddharth Sharma, Context Theory

Related questions

Is it better to have one general tool or several specific ones?

Several specific ones, up to the point of overlap. A general tool moves the decision inside the call, where it is invisible and cannot be constrained, while specific tools put the choice in a place you can see in the log and restrict by permission. The exception is where the specific tools would be so numerous that selection itself becomes the failure.

How do you tell whether a tool set is too large?

Look at the wrong calls. If the run log shows the agent trying two tools before finding the right one, or using a general tool for something specific, that is selection failure and it is a property of the set rather than of the run. Consistent first-choice accuracy is the signal that the set is the right size and adequately described.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide

2026 speed-to-lead benchmark · verified

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

What is specific to this page.

Evidence
Kind Claim Check it against
SoftwareTool selection is a decision made from a list of descriptions and degrades as the list grows and entries become harder to distinguish, so a short well-separated set produces higher first-choice accuracy than a broad one containing plausible near neighbours.Counting first-choice tool selection errors in run logs under a small and a large connected set for the same tasks.
WorkflowRemoving a tool the job does not require eliminates a class of unanticipated behaviour without any policy, prompt or review, which makes it the lowest-cost safety change available in most agent configurations.Listing the tools currently granted and identifying which appear in no step of the intended job.
ResponseOverlapping tools produce run-to-run inconsistency that presents as flakiness rather than as a configuration defect, so auditing a grown set for pairs whose applicability a reader could not distinguish is the corrective.Reading the tool descriptions in a connected set and marking any pair whose scope a competent reader could not separate.
Buying behaviourA specific tool places the selection decision in the run log where it can be inspected and permission-gated, while a general tool moves the same decision inside the call where it is invisible and cannot be constrained.Comparing what a run log records for a specific tool call against a general one performing the same work.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one