Context Theory Get your growth audit

Answer

How should an AI agent handle a production incident?

Reading, gathering and correlating. Not changing anything, because an incident is the worst moment for an unreviewed action.

Have it read, gather and correlate — logs, deploys, metrics, recent changes — and present what it found. Do not have it act. An incident is the moment when an unreviewed change is most likely and most damaging.

Incidents have a specific shape: the constraint is not skill, it is the time to assemble what happened from six systems that do not talk to each other. Logs in one place, deploy history in another, metrics in a third, the change that went out this morning in a fourth, an alert that fired an hour before anyone noticed. Gathering and lining these up is mechanical, tedious and slow, and it is where an agent contributes most.

So the useful output is a timeline with sources. What changed and when, what the error rate did and when, what deployed, what alerts fired, what the logs say at each transition. Presented as a sequence with the raw evidence attached, this compresses the first phase of an incident considerably, and it does so without anyone granting write access to anything.

The argument against letting it act is not caution in general; it is that incident conditions defeat the controls that make action safe. Nobody is reviewing carefully, the person nominally approving is doing three other things, the pressure to resolve is maximal, and the system state is by definition abnormal — which is exactly the situation an agent has least basis for reasoning about. Whatever permission policy exists in normal conditions is not the policy that will be applied at three in the morning.

There is a bounded exception that is worth naming honestly: a pre-agreed reversible action with a single obvious form, such as rolling back to the last known-good deployment. That is a decision made in advance, not during the incident, and it is safe because the action is fixed rather than composed. The distinction is between executing an agreed procedure and choosing what to do, and only the first is appropriate here.

Afterwards is where the second contribution sits, and it is undervalued. Writing the timeline into an incident record, correlating the symptoms with the change that caused them, drafting the summary, identifying which check would have caught it earlier. This is real work that usually gets done badly because the urgency has passed, and it is the part where an agent's tendency to be thorough about tedious things is genuinely useful.

One thing to prepare in advance: the agent needs read access to the systems that matter before the incident, not during. Configuring access under pressure is how permissions get granted broadly and never revisited, and an arrangement set up calmly is both narrower and available when it is needed.

The value during an incident is a timeline assembled in two minutes, not a fix applied in one.

Siddharth Sharma, Context Theory

Related questions

Can an agent diagnose the cause?

It can propose candidates and correlate evidence, which is useful, and its confidence should not be read as diagnosis. The characteristic failure is a plausible causal story built from a correlation, and during an incident that story is unusually persuasive because everyone wants one. Ask for the evidence for each candidate rather than for a conclusion.

What about routine alerts rather than incidents?

Those are a better fit for automation because they are predictable and low-stakes: gather the standard context, attach it to the alert, and tell whoever is on call what is already known. That removes most of the tedium of triage and none of the judgement, and it works precisely because the situation is not abnormal.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Firms that never responded to a web enquiry at all23%Category-wide
Average B2B first-response time42Category-wide

Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · hours · 1.25M inbound leads across 2,241 US firms · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowThe binding constraint in early incident response is assembling a timeline from systems that do not share a view — logs, deploys, metrics, alerts, recent changes — which is mechanical work an agent can compress substantially without any write access.Timing the manual assembly of a timeline for a past incident against the systems it required.
ConstraintIncident conditions defeat the controls that make agent action safe, because review is degraded, approval is distracted, resolution pressure is maximal and system state is abnormal, which is the situation the agent has least basis to reason about.Reviewing which approval steps were actually performed during the last incident.
ResponseA pre-agreed reversible action with a fixed form, such as rolling back to the last known-good deployment, differs from composed action because the decision was made in advance and the procedure is executed rather than chosen.Checking whether the candidate action is fully specified before the incident or assembled during it.
ProcurementRead access to incident-relevant systems must be configured before an incident, because access granted under pressure is granted broadly and is not subsequently narrowed.Reviewing when the current access grants for incident tooling were created and whether they were revisited.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one