
The useful question about AI return is not whether a vendor benchmark looks impressive. It is whether a specific workflow became better after the system went into use.
That requires a baseline. Without one, a faster-looking demonstration can hide extra review, new errors, support burden, or work that moved somewhere else.
Define the unit of work
Start by naming the work precisely.
"Improve customer service" is too broad. "Prepare a first response for routine delivery-status questions" is measurable. "Automate documents" is too broad. "Extract six fields from a recurring supplier form and route exceptions for review" is measurable.
For that unit of work, record:
- How often it happens
- How long it takes today
- Who performs and reviews it
- Which errors or exceptions occur
- Which systems and data it touches
- What happens when the workflow is unavailable
Measure value in more than one dimension
Time and direct cost
Track handling time, review time, rework, and the operating cost of the new system.
A simple starting formula is:
net time returned = manual time avoided - review time - exception handling - maintenance time
That is more honest than counting every generated draft as time saved.
Quality and risk
Faster work is not useful if correction or incident risk rises.
Measure the quality signals that fit the workflow:
- Field-level correction rate
- Factual corrections
- Incorrect routing
- Missed escalation
- Unauthorized actions
- Reopened cases
- Policy or data-boundary violations
High-impact errors should be tracked separately. Ten harmless formatting corrections are not equivalent to one incorrect financial commitment.
Capacity and service
Some systems create value by reducing queues or making a response available sooner, even when they do not remove a role.
Possible measures include:
- Queue age
- Time to first review
- Work completed within the service target
- Clean handoffs
- Unanswered requests
- Time available for higher-judgment work
Do not assign financial value to returned time unless the organization knows how that capacity will be used.
Operating burden
The pilot also creates new work. Include it.
- Model and infrastructure cost
- Integration maintenance
- Monitoring and evaluation
- Knowledge updates
- Incident investigation
- User support
- Recovery testing
A workflow can save handling time and still be a poor investment if it requires constant correction or specialist maintenance.
Use a scorecard without invented targets
Set targets from the current workflow and the business decision, not from a generic industry percentage.
| Measure | Baseline | Pilot evidence | Decision rule |
|---|---|---|---|
| Handling time | Measure current work | Measure the same unit with AI | Expand only if net time improves |
| Review and correction | Record current rework | Track reviewer changes | Keep human review until quality is stable |
| Exceptions | Count and classify | Compare volume and severity | Revise if high-impact exceptions increase |
| Operating cost | Record current tools and labor | Include model, support, and maintenance | Compare total cost, not model price alone |
| Service outcome | Choose one relevant signal | Measure during the pilot | Confirm the customer or operator outcome improves |
| Recovery | Document current fallback | Exercise pause and manual handling | Do not expand without a usable fallback |
The decision rule should be written before the pilot. Otherwise it is easy to move the goalposts after seeing the result.
Separate pilot evidence from projection
Early pilots use a small set of examples and close supervision. Production has more variation, changing data, unavailable systems, and users who behave differently from the test group.
Treat a pilot result as evidence about that pilot, not a guaranteed annual return. If the organization projects a larger benefit, state the assumptions: volume, adoption, quality, uptime, review effort, and maintenance cost.
Then revisit those assumptions after launch.
Know when to stop
An AI workflow may be technically possible and commercially wrong.
Stop or redesign when:
- Review takes as long as the original work
- High-impact errors remain difficult to detect
- The data boundary is wider than the benefit justifies
- A simpler rule or product solves the problem
- The workflow lacks an accountable owner
- Support cost grows faster than the value returned
Finding that during a bounded pilot is a good outcome. It prevents a weak workflow from becoming an expensive dependency.
The useful ROI question is simple: what changed, what did the change cost to operate, and does the evidence justify more authority or wider scope?