Answer
How do you tell whether a task is worth giving to AI?
Compare the time to check the output against the time to do the work yourself. If checking costs as much, skip it.
Compare what it costs to verify the output against what it costs to produce it yourself. Where verification is nearly as expensive as production, the task is not a candidate, however good the output looks.
The usual framing is whether the system can do the task. That is the wrong question, because the answer is increasingly yes and it does not tell you whether to. The right question is what the task costs you afterwards. Work that is fast to produce and slow to check is a bad trade at any level of quality, and work that is slow to produce and fast to check is a good one even when the output needs revision.
Three properties make verification cheap. The output has an external referent — a record it must match, a total it must reconcile to, a format that either parses or does not. The output is in your domain, so errors are visible to you at reading speed. And the output is bounded in length, because verification cost scales with volume in a way that production cost does not.
Three properties make it expensive. The output is a claim about something you would have to research to check. It is long, so checking means reading all of it carefully. Or it is a judgement — a strategy, an assessment, a recommendation — where there is no fact to compare against and the only test is whether you agree, which means you have done the thinking anyway.
This explains a pattern people find confusing: the tasks that feel most impressive to delegate are often the least worth delegating. A research summary on an unfamiliar subject is impressive and unverifiable. Reformatting a hundred records is unimpressive and checkable in seconds. The second saves time every week; the first mostly moves work from doing to worrying.
There is a second axis worth applying: how often does the task recur? A one-off worth an hour is rarely worth the effort of specifying properly for a machine, because most of the hour goes into stating the job. A weekly task worth twenty minutes is worth an afternoon of specification, because the specification is written once and amortises. Frequency is what turns a marginal case into an obvious one, and it is the reason recurring administrative work is the most reliably profitable category.
Finally, count the cost of a wrong answer that gets through, not just the cost of checking. Some tasks are cheap to verify and catastrophic to get wrong, and the correct treatment there is to delegate the production and keep the verification unconditional. Others are expensive to verify and harmless to get wrong, and the correct treatment is to delegate and not check at all, which is a legitimate decision that people rarely make explicitly and often make by accident.
Every task handed to a machine converts a production cost into a verification cost, and the trade is only worth making when the second is smaller.
Siddharth Sharma, Context Theory
Related questions
What about tasks nobody in the business can do at all?
Those are the highest-risk delegations, not the highest-value ones, because there is nobody to evaluate the output. Sometimes the right answer is still to proceed, with the understanding that you are buying a starting point rather than an answer. What does not work is treating the output as finished on the grounds that nobody can say otherwise.
Does this mean creative work is a poor fit?
It means judgement-heavy work has an expensive verification step, which is a different statement. Drafting is often worth delegating precisely because a bad draft is discarded quickly and a good one saves the blank page, and that is a cheap verification. The expensive case is the finished-looking artefact presented for approval, where the only way to evaluate it is to form the view yourself.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
| Odds of qualifying a lead — replying within the first hour vs after it | 7× | Category-wide |
2026 speed-to-lead benchmark · verified
Oldroyd, McElheran & Elkington, "The Short Life of Online Sales Leads", Harvard Business Review (March 2011) · 1.25M inbound leads across 2,241 US firms · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Delegating a task converts a production cost into a verification cost, so the decision turns on the ratio between the two rather than on whether the system is capable, and a task that is fast to produce and slow to check is a poor trade at any output quality. | Timing verification and production separately for a candidate task before delegating it. |
| Software | Verification is cheap where the output has an external referent to reconcile against, sits in the reader's own domain, and is bounded in length, because checking cost scales with volume while production cost does not. | Classifying a set of delegated tasks by these three properties and comparing recorded review times. |
| Buying behaviour | Recurrence rather than task size decides whether specification effort pays, because the cost of stating a job properly is incurred once and amortises across every subsequent run, which is why recurring administrative work is the most reliably profitable category. | Dividing the one-off specification time by the number of runs expected in a quarter. |
| Constraint | The cost of an error passing undetected is a separate axis from verification cost, producing two legitimate designs that are rarely chosen explicitly: unconditional checking for cheap-to-verify high-consequence work, and no checking at all for expensive-to-verify low-consequence work. | Stating the consequence of an undetected error for each delegated task and comparing it against the checking actually performed. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one