TL;DR. Most AI ROI exercises produce a defensible-looking number that nobody acts on. Six months later the number is embarrassing and the line item is harder to defend than it was before you measured anything. The job is to make a quarterly decision per team and workflow, using observable signals. This guide is that decision loop. A seat audit still has a place, but only as a secondary cut for flat-priced workspace plans.
Your CFO wants a number. A vendor handed you one. It says your AI program is producing a 3.7x productivity multiplier. You put it in a slide. The board nodded. Six months later, headcount has not moved, contractor spend has not moved, and the same CFO is asking why the line item is bigger than last year. You no longer want to talk about the 3.7x.
You have a decision problem dressed up as a measurement problem. The number on the slide was never going to drive an action. It was theater.
The good news is the decision loop is small, and you can run the first turn of it this week. The bad news is it requires giving up the comfort of a single number on a slide. Trade made willingly, you have a defensible AI program. Trade refused, you have another year of vendor charts that nobody acts on.
ROI frameworks hide the decision you need to make
The default executive measurement playbook for AI looks like this. A vendor-supplied productivity index based on a self-reported survey of “time saved per task.” A McKinsey-style ROI model with a 12-cell framework and three sensitivity scenarios. A quarterly business review slide that aggregates “AI productivity gains” into a single percentage. A perception survey of how much your team feels AI is helping.
Those tools all have a place as prompts for questions. None belongs in the board number. Self-report is especially weak as a finish line. In a 2025 randomized trial of 16 experienced open-source developers working on 246 real tasks, developers forecast a 24% speedup and later believed they had gained 20%; measured task time was 19% slower with AI.1 That study is not a universal estimate of AI’s effect. It is a useful warning: felt speed and measured work can point in opposite directions.
The deeper problem is that none of these outputs has a decision attached. Whether the index reads 3.7x or 4.1x, nothing about the program changes. Track only numbers that change a decision. A measurement that does not change a decision is a ritual.
MIT NANDA’s 2025 preliminary report found that 95% of the organizations it studied showed no measurable P&L return from generative AI.2 The figure is a warning about the pilot treadmill, not a reason to give up on the work. A pilot has to become an owned production workflow before it counts. Report production, not pilots. Why most pilots stall before that point is a separate problem from whether the model can answer a prompt.
The CFO eventually figures this out. Usually around the second or third budget cycle, after enough time has passed that the promised gains should be visible somewhere in the P&L and they are not. At that point the ROI slide stops being a defense of the line item and starts being evidence that the line item should be cut. The shorter case against the ROI framework is the one to send before that meeting.
The cost side changed shape too
There is a second reason the old number is failing you. On a consumption-priced plan, a team whose token spend doubled while its cycle time halved is the best thing in your portfolio. It is a win to fund, not a cost to cap.
The unit of measurement moves with it, from seat count to cost per outcome. The question is no longer “are we paying for idle seats.” It is “what did this consumption buy.” Spend rising against flat workflow signals is the problem. Spend rising against cycle times that fell by half is the program working out loud. Throttle the first. Fund the second.
The efficiency levers moved too. A cache read can cost one tenth of a normal input token, which is a 90% reduction on the cached input only. Cache writes cost more, so the saving arrives only after repeated context has earned back that premium. Batch processing cuts both input and output token prices by 50% for work that can wait, and it can be combined with caching.3 One person who reads the usage dashboard monthly and understands those two levers can take more off a metered bill than a quarter of seat-culling ever did. The seat audit did not disappear. It dropped from the headline to the footnote.
AI returns concentrate in people and workflows
An averaged ROI number cannot tell you anything useful because AI returns are not distributed like SaaS returns. A CRM produces a small uniform gain across every seat. AI produces a step change in a small number of people and roughly nothing in the rest.
The averages still give you a useful warning about what they cannot prove. Bick, Blandin, and Deming’s February 2025 St. Louis Fed analysis found that workers who had used generative AI in the previous week self-reported saving 5.4% of work hours, about 2.2 hours in a 40-hour week.4 That is evidence about reported experience among active users, not an outcome your CFO should book. OpenAI’s 2025 enterprise report found its 95th-percentile workers sent about six times as many messages as the median employee.5 That is an engagement gap, not a productivity gap. Both figures explain why the average hides where to look.
The head of the distribution used to be the people with the temperament to wring leverage out of a chat tab. It now includes a second type: the person who designs a recurring workflow once and schedules it, creating capacity for a whole team that never changed a habit. Disposition and design both. Averaging either into the tail erases the only thing worth measuring.
The returns are also lagged. A champion in operations who automates the monthly close in February does not show up in the budget until the next hiring cycle decides not to backfill the role she was about to be promoted out of. A marketing team that ships campaigns without an agency does not show up until the agency contract comes up for renewal nine months later. Quarterly ROI snapshots miss this entirely. They take a picture of a process whose financial signal arrives a year late.
The right measurement layer is therefore not “what is the program ROI.” It is three smaller, more boring questions, asked at three different cadences.
The three layers that drive decisions
There are three measurement layers. Each runs on its own clock and produces its own decision. None produce a single number. All produce something you can act on.
Workspace seats, monthly. Decision: reallocate a flat-priced seat or leave it alone. A fixed-price workspace seat is a scarce allocation. Review it against two questions: is there meaningful consumption, and can the manager point to observable work change in the last thirty days? Meaningful consumption is work the vendor can show, not logged-in time. Observable change means something shipped faster, cheaper, or at all that would not have happened otherwise. A seat with neither signal for sixty days moves to the waitlist. The rule belongs only on a flat-priced workspace SKU. On a metered plan, low activity costs little beyond the base fee; the relevant measure is team consumption against team outcomes. High consumption with no work change is where that review starts. The mechanics of the spend audit are in Evaluating Spend.
Workflow-level, quarterly. Decision: invest more, leave alone, or formalize. Once a quarter, list the five to ten recurring workflows in your org where AI is most likely to be landing. Month-end close. Contract redlining. Campaign briefing. Pipeline research. First-pass code review. For each one, ask whether cycle time, cost, or staffing has moved in the last ninety days. The threshold is half or more. If it has, that workflow gets more investment. Better tooling, dedicated time, and a champion with enough authority to carry the workflow into the rest of the business. If it has not, leave it alone for another quarter. If the workflow has been transformed and now runs on a personal account or a hand-built script, formalize it. Pay for the tool. Document the prompt. Make it survivable when the person leaves.
Add one more question to the quarterly pass: is a recurring report a person used to assemble now produced by a scheduled automation? That is persistent capacity, and it is load-bearing work you have to inventory and own. The Q3 2026 briefing covers why the inventory is now a risk issue as well as a returns issue.
Org-level, annually. Decision: rebalance the workforce plan. Once a year, before the headcount planning cycle, look for the workforce-plan changes AI helped create. A contractor renewal that came in lower. A backfill that got absorbed. A team that can hold its output with fewer planned additions. These are real financial signals, but they are not a rule that a growing business should stop hiring. Growth can justify more people even when each person has more capacity. Compare the plan against demand and output, then ask whether AI changed the capacity behind it. The answer appears in the diff between this year’s workforce plan and last year’s, not on a vendor dashboard.
Three layers, three clocks, three decisions. The aggregate productivity number is gone. In its place is a program you can steer.
Four observable changes show whether the program is landing
When an AI program is working at the org level, you will see at least three of the following four changes inside twelve months. If you see fewer than three, the program is not landing, regardless of what your dashboard says.
-
A recurring report or process now takes less than half the time it used to. Month-end close. Board prep. Pipeline review. Quarterly business review deck. Something that used to be a multi-day exercise is now a half-day exercise, and the person doing it can name the change.
-
A workforce-plan request changed. A team that was scoped to grow asks for less because its output held. A backfill does not get filed. A contractor scope comes in smaller. The work and demand still need to be visible beside the plan, or you are only measuring a paused requisition.
-
A vendor or contractor line item dropped. Agency spend going down. Outside counsel spend on routine matters going down. A SaaS tool getting canceled because the workflow it supported is now done inside a chat window. The vendor relationship manager will notice this before your finance team does.
-
Reusable work is visible and shared. A prompt template that started in one team is now in three. A script someone wrote on a weekend now has an owner. A scheduled automation is in the inventory instead of hidden in a personal account. This is the strongest leading indicator. It means capacity is no longer trapped inside one person.
That is the whole heuristic. Three of four, inside a year. If you have it, the program is working and you should fund the layers that produced it. If you do not, the program is not working and another training cohort will not fix it.
Flat-priced seats should move when use and outcomes are absent
Killing idle seats used to be the headline efficiency move. It now applies to a narrower case: a flat-priced workspace plan where access itself is the scarce resource. On a metered bill, the money is in consumption and the plumbing behind it. An unused seat carries little marginal cost beyond its base fee.
Write the rule before you need it. A flat-priced workspace seat moves after sixty days when the usage report shows no meaningful consumption and the manager cannot name an observable work change. Hours in the tool are evidence at most. They are not the rule. There are no exceptions for title, tenure, or personal preference. The open seat belongs with the next person who can name a recurring job for it.
The waitlist is the second half of the rule. There should always be people who have asked for a seat and not been given one. When a seat opens, it goes to the top of that list. Seats migrate toward people who use them.
The cost is consumption that buys no work change, plus the capacity you did not give to someone who would use it. That is a better unit than a made-up per-seat dollar figure, and it is the one the next budget conversation can defend.
The CFO report needs consumption and an agent inventory
The CFO does not want a productivity multiplier. The CFO wants a defensible paragraph for the budget conversation. Here is the format that holds up.
AI program spend this quarter was $X, against $W last quarter. Consumption concentrated in teams A, B, and C, where it bought cycle-time reductions of 50% or more in core workflows. Our agent inventory contains N scheduled automations we can see and M reported or suspected automations still awaiting confirmation and an owner. Two workforce-plan requests changed this cycle, representing approximately $Y in avoided cost. Outside vendor spend in [category] is down $Z year over year, attributable to internal AI capability. We reduced the high-volume jobs with caching and batching, and reallocated unused flat-priced workspace seats to the waitlist. Next quarter we are funding the teams where spend is buying the most cycle time and formalizing workflow D.
That is the entire report. Consumption by team, the work it bought, the visible and suspected automation count, workforce capacity, and decisions made. The automation pair belongs here because a recurring process you cannot see or assign is not capacity you can safely count on. What your board will ask about AI covers the ownership question from the board side.
If you cannot fill in three of the four signals in that paragraph, the program is not producing returns and no measurement framework is going to manufacture them. That is useful information. It tells you where to stop measuring and start changing the work.
Something to carry
The report to carry into the next conversation starts with the largest vendor’s consumption report grouped by team. Put the top three teams next to the recurring workflow each is paying for, its cycle-time or cost change, and the scheduled automations you can name with an owner. Beside that, record the scheduled automations managers say exist but you cannot yet see. That is the first pass at a returns report.
Below it, attach the secondary cut: flat-priced workspace seats with no meaningful consumption. For each, the manager answers two questions. Has this person produced observable work change with the tool in the last thirty days? Is there someone on the team who has a recurring job for the seat if it opens? The answers create a reallocation list without confusing logged hours for return.
The next month gives you a better comparison: team consumption against the work it bought, a visible automation count against the suspected count, and a small group of fixed-price seats moving toward active work. That is enough to make the next decision.
Footnotes
-
METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, July 10, 2025. The randomized study involved 16 experienced developers and 246 tasks in mature open-source repositories. It does not generalize to every role or tool. Last verified August 8, 2026. ↩
-
MIT NANDA, The GenAI Divide: State of AI in Business 2025, preliminary findings, July 2025. The report analyzed 300 public implementations and describes 95% of organizations as getting zero measurable return. Treat the headline as the report’s finding, not a universal failure rate. Last verified August 8, 2026. ↩
-
Anthropic, “Pricing”. Cache hits are priced at 0.1x standard input tokens; five-minute cache writes are 1.25x and one-hour writes 2x. The Batch API applies a 50% discount to input and output tokens, and Anthropic says batch and caching discounts can combine. Last verified August 8, 2026. ↩
-
Alexander Bick, Adam Blandin, and David Deming, “The Impact of Generative AI on Work Productivity”, Federal Reserve Bank of St. Louis, February 27, 2025. Based on the November 2024 Real-Time Population Survey wave; the 5.4% and 2.2-hour figures are self-reported by workers who used generative AI in the previous week. Last verified August 8, 2026. ↩
-
OpenAI, “The State of Enterprise AI”, December 2025. Its 95th-percentile workers sent 6x the messages of the median employee. Message volume is engagement data, not a productivity measure. Last verified August 8, 2026. ↩