AI budgets are growing, and so is pressure to show results. Yet many organizations struggle to answer a basic question: is this actually paying off? The problem is rarely the technology. It is that nobody defined what success looks like before the project started.
Updated September 2026: we added two studies that show how review time and worker experience change the returns from AI.
Start with a specific process
“Use AI to be more productive” cannot be measured. “Reduce time to draft first responses to support tickets” can. Pick a process with:
- A clear start and end
- Enough volume to measure
- Existing data, or data you can easily collect
Set a baseline first
Before introducing AI, measure how the process performs today for two to four weeks. Useful baseline metrics include:
- Time per task
- Volume handled per person
- Error or rework rate
- Customer satisfaction or quality score
- Cost per unit of output
Without a baseline, any improvement is a guess.
Count the full cost
| Cost | Often forgotten? |
|---|---|
| Licenses and usage fees | No |
| Setup and integration | Sometimes |
| Training time | Yes |
| Time spent reviewing and correcting AI output | Very often |
| Governance, security and policy work | Yes |
| Ongoing maintenance of prompts and content | Yes |
Review time is the one most often missed, and it can outweigh the gains: in a randomized trial by METR in 2025, experienced developers took 19% longer on tasks when using AI tools, even though they believed they had been faster. If a draft takes five minutes to generate and twenty minutes to fix, the net benefit may be small or negative.
Measure outcomes, not usage
Number of prompts sent or active users tells you people are trying the tool, not that it helps. Focus on the process metrics you baselined, plus quality. A faster process that produces more errors is not a win.
Simple ROI formula Monthly value = (time saved per task x tasks per month x loaded hourly cost) + measurable quality gains. ROI = (value minus total cost) divided by total cost. Keep assumptions explicit so others can challenge them.
Include the qualitative picture
Some benefits are hard to quantify: less tedious work, faster onboarding for new staff, better consistency. The onboarding effect can be large: in a study of 5,179 customer support agents, an AI assistant raised productivity 14% on average and 34% for novice workers. Collect short structured feedback from users. Treat it as supporting evidence rather than the main case.
Decide in advance when to scale or stop
Set thresholds before the pilot: for example, scale if time per task falls by at least 20 percent with no drop in quality after eight weeks. Pre-committing prevents enthusiasm, or skepticism, from deciding the outcome.
Common pitfalls
- Measuring during the novelty period only
- Comparing the best AI-assisted users with average non-users
- Ignoring hidden rework
- Spreading a pilot across too many use cases to measure any of them well
Our guides to getting started with AI and writing an AI policy cover the groundwork that makes measurement easier.
Making the numbers count
AI return on investment is measurable if you pick a specific process, record a baseline, count every cost including review time and judge outcomes rather than usage. That discipline turns AI from an act of faith into a business decision.
Sources
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, July 2025
- Generative AI at Work, National Bureau of Economic Research, 2023



