Answer
How do you tell whether an AI project actually worked?
Whether the thing it was supposed to change moved, and whether it is still running six months later. Both questions get avoided.
The measure it was supposed to move actually moved, and the thing is still running once nobody is paying attention. Delivery is not success and enthusiasm is not evidence, and both are what usually gets reported.
Two questions decide this and both are avoidable, which is why they are usually avoided. Did the number the project existed to change actually change? And is it still working after the people who built it moved on? Reporting delivery instead — it was built, it works, people like it — answers neither, and a project can pass all three of those and have changed nothing.
The first question requires a baseline, which means it was either measured before the project or it cannot be answered now. This is the single most common reason these assessments end in argument: nobody knows what the response time, the error rate or the handling time was beforehand, so the comparison is between a measurement and a recollection, and recollections adjust to accommodate effort already spent.
The number should be the outcome rather than the activity. Items processed, messages sent and time in the tool all measure that the thing ran. What the project existed to change is downstream: how quickly enquiries got a real reply, how many fell through, how much editing the output needed, how long the process took end to end. Choosing the outcome measure at the start is uncomfortable precisely because it can fail, which is what makes it the useful one.
The second question is the one nobody asks, and it is the more discriminating. A large share of these projects work at launch and are not running a year later, having stopped when something upstream changed and nobody noticed. Checking at six months costs a moment and answers whether what was built was maintainable, which is a different property from whether it worked and is the one that determines the return.
Two secondary measures are worth capturing. Whether the human time on the process actually fell, measured rather than assumed, since automations frequently relocate work rather than removing it. And what the total cost was including the build, the maintenance and the exception handling, against the saving. Both are unglamorous and both prevent the common outcome where a project is considered a success on the basis that the software functions.
Finally, count what was learned as part of the return. A project that did not move its number and established that the process was more varied than anyone thought, or that the data everyone relied on is unreliable, has produced something worth having. That is a legitimate outcome and it should be reported as what it is rather than as a success, because reporting it as a success is how the same project gets attempted again.
Ask in month six, not month one, because month one measures the attention and month six measures the system.
Siddharth Sharma, Context Theory
Related questions
What if there was no baseline?
Establish one now for the current state and accept that the comparison with before is unavailable. That is an honest position and it makes the next assessment possible, which is more useful than an argument about what things used to be like. Where records exist, a historical period can sometimes be reconstructed, and it is worth an hour to try.
How long should you wait before judging?
Long enough for the attention to have gone, which is usually a few months rather than weeks. Early measurement captures a period when people are watching, correcting and interested, and none of those conditions persists. The number that matters is the one from an ordinary month.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
| Close rate — response under 5 minutes vs over 24 hours | 32% vs 12% | Category-wide |
2026 speed-to-lead benchmark · verified
Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Delivery, functionality and user satisfaction can all be positive while the outcome the project targeted is unchanged, so reporting them answers a different question from whether the project worked. | Comparing the reported outcome of a completed project against movement in the measure it was intended to change. |
| Constraint | Without a baseline measured before the project, the comparison is between a measurement and a recollection, and recollections adjust to accommodate effort already expended. | Asking two people who were involved to state what the pre-project figure was, independently. |
| Response | A substantial share of these projects function at launch and are not running a year later, having stopped when something upstream changed, which makes survival at six months a more discriminating measure than launch success. | Checking which projects delivered in the last two years are still operating. |
| Buying behaviour | Early measurement captures a period of active attention, correction and interest that does not persist, so the figure that represents the system rather than the attention comes from an ordinary month some time later. | Comparing the measured outcome in the first month against the sixth. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one