01. Why ROI Is Hard to Measure
The second common failure is measuring the wrong thing entirely: volume of output instead of business outcome. A dashboard showing thousands of messages generated says nothing about whether those messages replaced hours of work, improved accuracy, or actually shipped.
02. The Formula That Holds Up
- Hours returned — time no longer spent by a person on the task, measured against the pre-agent baseline.
- Error rate change — quality of outcome compared to the human-only process it replaces or augments.
- Cycle time — how long a task takes from start to finish, end to end.
- Cost per completed task — total cost, including engineering amortised over volume, divided by tasks actually finished.
Track more than one if you want a fuller picture, but report against the one you named before launch. Switching the headline metric after the fact is the fastest way to lose credibility with whoever's holding the budget.
03. Building a Baseline First
- Time a representative sample of the task as it's done today.
- Record the current error or rework rate, not an estimate of it.
- Note the current cost per unit — a support ticket, a processed invoice, a qualified lead.
This takes days, not months, and it's the single highest-leverage step in the entire project — everything downstream depends on having a real number to compare against.
04. Common ROI Killers
- No baseline measured before launch, so the "after" number has nothing credible to compare against.
- Scope creep during the build that inflates engineering cost past what the workflow's volume justifies.
- Low trust in the output, keeping a human reviewing every result and eliminating the time savings the case was built on.
- Picking a workflow with too little volume for the fixed engineering cost to pay back in a reasonable window.
Our AI agent cost guide covers the spending side of this equation in detail.
05. A Simple ROI Worksheet
06. Frequently Asked
What's a good metric for AI agent ROI?
Hours returned, error rate change, and cost per completed task, each measured against a baseline recorded before the agent launched. 'Messages generated' or 'tasks handled' without a quality or cost comparison isn't ROI, it's activity — and activity metrics are exactly what gets a project cancelled at renewal when someone asks what it actually saved.
How long before an AI agent shows ROI?
A narrow, well-scoped pilot can show a measurable signal within the first few weeks, since the baseline and the comparison are both available quickly at that scale. Broader deployments take longer, not because the agent is slower to help but because the change management and trust-building around wider adoption takes real time.
What if the ROI case doesn't work out?
That's a legitimate outcome, and worth planning for. If a rigorous baseline shows the agent isn't moving the metric you picked, the fix is usually scope — narrow the workflow further, fix upstream data quality, or retire the pilot and pick a better-fitted candidate. Treat 'no' as useful information from a real experiment, not a reason to hide the result.
Cloudz Computing scopes every engagement around a measured baseline and a named metric, before a single line of the system gets built.
Request a private consultation