Measuring Returns

Updated

TL;DR. Most AI ROI exercises produce a defensible-looking number that nobody acts on. Six months later the number is embarrassing and the line item is harder to defend than before you measured anything. The job is to make a quarterly decision per team and workflow, using observable signals. The quarterly decision is the measurement. A seat audit still has a place, but only as a secondary cut for flat-priced workspace plans.

Your CFO wants a number. A vendor handed you one. It says your AI program is producing a 3.7x productivity multiplier. You put it in a slide. The board nodded. Six months later, headcount has not moved, contractor spend has not moved, and the same CFO is asking why the line item is bigger than last year. You no longer want to talk about the 3.7x.

You have a decision problem dressed up as a measurement problem. The number on the slide was never going to drive an action. It was theater.

The decision loop is small, and you can run the first turn this week. It requires giving up the comfort of a single number on a slide. In exchange, you get an AI program you can defend with decisions rather than vendor charts.

ROI frameworks hide the decision you need to make

The default executive measurement playbook for AI looks like this. A vendor-supplied productivity index based on a self-reported survey of “time saved per task.” A McKinsey-style ROI model with a 12-cell framework and three sensitivity scenarios. A quarterly business review slide that aggregates “AI productivity gains” into a single percentage. A perception survey of how much your team feels AI is helping.

Those tools can prompt useful questions. None belongs in the board number. Self-report is especially weak as a finish line. In a 2025 randomized trial of 16 experienced open-source developers working on 246 real tasks, developers forecast a 24% speedup and later believed they had gained 20%; measured task time was 19% slower with AI.1 That study does not estimate AI’s effect everywhere. It is a useful warning: felt speed and measured work can point in opposite directions.

The deeper problem is that none of these outputs has a decision attached. Whether the index reads 3.7x or 4.1x, nothing about the program changes. Track only numbers that change a decision. A measurement that does not change a decision is a ritual.

MIT NANDA’s 2025 preliminary report found that 95% of the organizations it studied showed no measurable P&L return from generative AI.2 The figure warns about the pilot treadmill. Only an owned production workflow counts. Why most pilots stall before that point is a separate problem from whether the model can answer a prompt.

The CFO eventually figures this out. Usually around the second or third budget cycle, after enough time has passed that the promised gains should be visible somewhere in the P&L and they are not. At that point the ROI slide stops being a defense of the line item and starts being evidence that the line item should be cut. The shorter case against the ROI framework is the one to send before that meeting.

The cost side changed shape too

There is a second reason the old number is failing you. On a consumption-priced plan, a team whose token spend doubled while its cycle time halved is the best thing in your portfolio. It deserves more funding.

The unit of measurement moves with it, from seat count to cost per outcome. Ask what the consumption bought. Spend rising against flat workflow signals is the problem. Spend rising against cycle times that fell by half is the program working out loud. The first calls for a throttle. The second deserves funding.

The efficiency levers moved too. A cache read can cost one tenth of a normal input token, which is a 90% reduction on the cached input only. Cache writes cost more, so the saving arrives only after repeated context has earned back that premium. Batch processing cuts both input and output token prices by 50% for work that can wait, and it can be combined with caching.3 One person who reads the usage dashboard monthly and understands those two levers can take more off a metered bill than a quarter of seat-culling ever did. The seat audit dropped from the headline to the footnote.

AI returns concentrate in people and workflows

An averaged ROI number cannot tell you anything useful because AI returns are not distributed like SaaS returns. A CRM produces a small uniform gain across every seat. AI produces a step change in a small number of people and little in the rest.

The averages still give you a useful warning about what they cannot prove. Bick, Blandin, and Deming’s February 2025 St. Louis Fed analysis found that workers who had used generative AI in the previous week self-reported saving 5.4% of work hours, about 2.2 hours in a 40-hour week.4 That is evidence about reported experience among active users, not an outcome your CFO should book. OpenAI’s 2025 enterprise report found its 95th-percentile workers sent about six times as many messages as the median employee.5 That is an engagement gap, not a productivity gap. Both figures explain why the average hides where to look.

The head of the distribution used to be people with the temperament to wring leverage out of a chat tab. It now includes a second type: the person who designs a recurring workflow once and schedules it, creating capacity for a whole team that never changed a habit. Both disposition and design matter. Averaging either into the tail erases the thing worth measuring.

The returns are also lagged. A champion in operations who automates the monthly close in February does not show up in the budget until the next hiring cycle decides not to backfill the role she was about to be promoted out of. A marketing team that ships campaigns without an agency does not show up until the agency contract comes up for renewal nine months later. Quarterly ROI snapshots miss this entirely. They take a picture of a process whose financial signal arrives a year late.

Use three smaller, more boring questions at three different cadences instead of asking for program ROI.

The three layers that drive decisions

There are three measurement layers. Each runs on its own clock and produces its own decision. None produce a single number. All produce something you can act on.

Workspace seats, monthly. Decision: reallocate a flat-priced seat or leave it alone. A fixed-price workspace seat is a scarce allocation. Review it against two questions: is there meaningful consumption, and can the manager point to observable work change in the last thirty days? Meaningful consumption is work the vendor can show, not logged-in time. Observable change means something shipped faster, cheaper, or at all that would not have happened otherwise. A seat with neither signal for sixty days moves to the waitlist. The rule belongs only on a flat-priced workspace SKU. On a metered plan, low activity costs little beyond the base fee; the relevant measure is team consumption against team outcomes. High consumption with no work change is where that review starts. The mechanics of the spend audit are in Evaluating Spend.

Workflow-level, quarterly. Decision: invest more, leave alone, or formalize. Once a quarter, list the five to ten recurring workflows in your org where AI is most likely to be landing. Month-end close. Contract redlining. Campaign briefing. Pipeline research. First-pass code review. For each one, ask whether cycle time, cost, or staffing has moved in the last ninety days. The threshold is half or more. If it has, that workflow gets more investment. Better tooling, dedicated time, and a champion with enough authority to carry the workflow into the rest of the business. If it has not, leave it alone for another quarter. If the workflow has been transformed and now runs on a personal account or a hand-built script, formalize it. Pay for the tool. Document the prompt. Make it survivable when the person leaves.

Add one more question to the quarterly pass: is a recurring report a person used to assemble now produced by a scheduled automation? That is persistent capacity, and work other people depend on needs an inventory and an owner. The Q3 2026 briefing covers why the inventory is now a risk issue as well as a returns issue.

Org-level, annually. Decision: rebalance the workforce plan. Once a year, before the headcount planning cycle, look for the workforce-plan changes AI helped create. A contractor renewal that came in lower. A backfill that got absorbed. A team that can hold its output with fewer planned additions. These are real financial signals, but they are not a rule that a growing business should stop hiring. Growth can justify more people even when each person has more capacity. Compare the plan against demand and output, then ask whether AI changed the capacity behind it. The answer appears in the diff between this year’s workforce plan and last year’s, not on a vendor dashboard.

Three layers, three clocks, three decisions. The aggregate productivity number is gone. In its place is a program you can steer.

Four observable changes show whether the program is landing

When an AI program is working at the org level, you will see at least three of the following four changes inside twelve months. If you see fewer than three, the program is not landing, regardless of what your dashboard says.

  1. A recurring report or process now takes less than half the time it used to. Month-end close. Board prep. Pipeline review. Quarterly business review deck. Something that used to be a multi-day exercise is now a half-day exercise, and the person doing it can name the change.

  2. A workforce-plan request changed. A team that was scoped to grow asks for less because its output held. A backfill does not get filed. A contractor scope comes in smaller. The work and demand still need to be visible beside the plan, or you are only measuring a paused requisition.

  3. A vendor or contractor line item dropped. Agency spend going down. Outside counsel spend on routine matters going down. A SaaS tool getting canceled because the workflow it supported is now done inside a chat window. The vendor relationship manager will notice this before your finance team does.

  4. Reusable work is visible and shared. A prompt template that started in one team is now in three. A script someone wrote on a weekend now has an owner. A scheduled automation is in the inventory instead of hidden in a personal account. This is the strongest leading indicator. It means capacity is no longer trapped inside one person.

That is the whole heuristic. Three of four, inside a year. If you have it, the program is working and you should fund the layers that produced it. If you do not, the program is not working and another training cohort will not fix it.

Flat-priced seats should move when use and outcomes are absent

Killing idle seats used to be the headline efficiency move. It now applies to a narrower case: a flat-priced workspace plan where access itself is the scarce resource. On a metered bill, the money is in consumption and the plumbing behind it. An unused seat carries little marginal cost beyond its base fee.

The rule belongs in writing before you need it. A flat-priced workspace seat moves after sixty days when the usage report shows no meaningful consumption and the manager cannot name an observable work change. Hours in the tool are evidence at most. They do not decide the outcome. There are no exceptions for title, tenure, or personal preference. The open seat belongs with the next person who can name a recurring job for it.

The waitlist is the second half of the rule. There should always be people who have asked for a seat and not been given one. When a seat opens, it goes to the top of that list. Seats migrate toward people who use them.

The cost is consumption that buys no work change, plus the capacity you did not give to someone who would use it. That is a better unit than a made-up per-seat dollar figure, and it is the one the next budget conversation can defend.

The CFO report needs consumption and an agent inventory

For the budget conversation, the CFO needs a defensible paragraph rather than a productivity multiplier. Here is the format that holds up.

AI program spend this quarter was $X, against $W last quarter. Consumption concentrated in teams A, B, and C, where it bought cycle-time reductions of 50% or more in core workflows. Our agent inventory contains N scheduled automations we can see and M reported or suspected automations still awaiting confirmation and an owner. Two workforce-plan requests changed this cycle, representing approximately $Y in avoided cost. Outside vendor spend in [category] is down $Z year over year, attributable to internal AI capability. We reduced the high-volume jobs with caching and batching, and reallocated unused flat-priced workspace seats to the waitlist. Next quarter we are funding the teams where spend is buying the most cycle time and formalizing workflow D.

The report covers consumption by team, the work it bought, the visible and suspected automation count, workforce capacity, and decisions made. The automation pair belongs here because a recurring process you cannot see or assign cannot be counted as capacity. What your board will ask about AI covers the ownership question from the board side.

If you cannot fill in three of the four signals in that paragraph, the program is not producing returns and no measurement framework is going to manufacture them. That is useful information. It tells you where to stop measuring and start changing the work.

A returns report begins with consumption by team

The report to carry into the next conversation starts with the largest vendor’s consumption report grouped by team. Put the top three teams next to the recurring workflow each is paying for, its cycle-time or cost change, and the scheduled automations you can name with an owner. Beside that, record the scheduled automations managers say exist but you cannot yet see. That is the first pass at a returns report.

Below it, attach the secondary cut: flat-priced workspace seats with no meaningful consumption. For each, the manager answers two questions. Has this person produced observable work change with the tool in the last thirty days? Is there someone on the team who has a recurring job for the seat if it opens? The answers create a reallocation list without confusing logged hours for return.

The next month gives you a better comparison: team consumption against the work it bought, a visible automation count against the suspected count, and a small group of fixed-price seats moving toward active work. That is enough to make the next decision.

Footnotes

  1. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, July 10, 2025. The randomized study involved 16 experienced developers and 246 tasks in mature open-source repositories. It does not generalize to every role or tool. Last verified August 8, 2026.

  2. MIT NANDA, The GenAI Divide: State of AI in Business 2025, preliminary findings, July 2025. The report analyzed 300 public implementations and describes 95% of organizations as getting zero measurable return. Treat the headline as the report’s finding, not a universal failure rate. Last verified August 8, 2026.

  3. Anthropic, “Pricing”. Cache hits are priced at 0.1x standard input tokens; five-minute cache writes are 1.25x and one-hour writes 2x. The Batch API applies a 50% discount to input and output tokens, and Anthropic says batch and caching discounts can combine. Last verified August 8, 2026.

  4. Alexander Bick, Adam Blandin, and David Deming, “The Impact of Generative AI on Work Productivity”, Federal Reserve Bank of St. Louis, February 27, 2025. Based on the November 2024 Real-Time Population Survey wave; the 5.4% and 2.2-hour figures are self-reported by workers who used generative AI in the previous week. Last verified August 8, 2026.

  5. OpenAI, “The State of Enterprise AI”, December 2025. Its 95th-percentile workers sent 6x the messages of the median employee. Message volume is engagement data, not a productivity measure. Last verified August 8, 2026.