Evaluating Spend

Updated

TL;DR. Most AI budgets are wrong in the same direction. A flat per-seat license sits in the inboxes of people who never open it, while the three people producing real leverage put API spend on a personal card. Seats still anchor the price. Metering sits on top, where agent work and APIs land. A five-minute first pass exposes both allocation problems.

Your finance team is asking what the AI line item is producing. A vendor wants forty-five minutes to walk you through a higher tier. Three department heads want to expense Claude Team seats independently. A board member sent you a podcast. Somewhere in your expense system, a senior engineer is reimbursing a Claude Pro or Max account because the approved tool is worse than the free tier of the unapproved one.

The allocation determines whether the dollar total produces anything useful.

Your AI line item hides the allocation

The default AI budget reads like this. A bundled Copilot SKU rolled into your existing Microsoft 365 contract at $30 per seat per month. A pilot of ChatGPT Business at $20 per user per month on an annual plan, or $25 month to month.1 A handful of Claude Team seats that your data lead bought after a conference. Maybe an API budget if you have engineers, usually buried in the cloud line and untracked. Add it up and you’re spending somewhere between $40 and $150 per employee per month on AI, depending on how generous your seat distribution is and how much your engineers are quietly burning on tokens.

That number, by itself, tells you almost nothing. Start with whom the money reaches and what work it funds. A $90 per-employee average can be a leveraged spend if it is concentrated where the leverage lives. The same $90 can be pure shelfware if it is sprayed evenly across a workforce that mostly does not open the tool.

Most AI budgets are the second case. Your CFO is asking the wrong question because you handed them the wrong instrument. The budget is in the wrong shape, and the audit makes that visible.

Three budget patterns hollow out the line item

Almost every AI budget I’ve seen is some combination of these three patterns. Each looks defensible in isolation. Together they produce a line item with no leverage to point at.

The bundled-with-the-suite mistake. You added Copilot to your E5 license because the rep made it easy and procurement preferred a single contract. Microsoft reported more than 20 million paid M365 Copilot seats in Q3 FY26. Against its 450 million-plus commercial M365 base, that is roughly 4.4% penetration.2 When users at the same company get a free choice between Copilot, ChatGPT, and Gemini, they pick ChatGPT seventy-six percent of the time. Recon Analytics recorded negative Copilot accuracy NPS at each of its published panel readings: -3.5 in July 2025, -24.1 in September, and -19.8 in January 2026.3 You bought it because it was easy. Your people don’t open it because it isn’t the tool they would choose.

The $30 add-on has not moved, but the underlying suites did on July 1. At list price, M365 Copilot now puts an E3 seat around $69 all-in and an E5 seat around $90. For organizations with up to 300 seats, Copilot Business is $18 on promotion or $21 list through September 30, 2026.4 Audit the all-in number, not the add-on alone. A $30 decision can look modest only because most of its cost is already hiding in another line.

The flat per-seat democracy mistake. Everyone gets a seat. Marketing, finance, legal, ops, the loading dock, the regional VP who hasn’t opened a chat interface in his life. The logic is fairness, or future-proofing, or “we want to give everyone the chance to learn.” The result is a usage curve where a top few percent of seats produce most of the value and a third of the pool never logs in. You’re paying full price for the median, and the median is zero.

The training-budget-as-AI-budget mistake. You allocated $40,000 to AI training this year. A vendor delivered six lunch-and-learns. Attendance was strong. Usage didn’t move. The engagement gap between power users and everyone else, better than tenfold by volume in the telemetry, isn’t a knowledge gap, it’s a disposition gap, and a disposition gap doesn’t close with curriculum. You spent your AI budget on the symptom and got the symptom back.

These three patterns share a shape. They optimize for procurement convenience, for political fairness, and for the appearance of action. None of them are anchored to where the leverage actually lives in your org.

The bill has separate seat and metered costs

The mid-market plans most leaders buy still charge a flat seat fee. Claude Team, ChatGPT Business, M365 Copilot, and Google Workspace all include defined usage in the seat price.5 That makes seat distribution matter. An idle license costs the same as one someone opens every day, so the deployment has to earn renewal one seat at a time.

Metered spend sits on top. It catches APIs, agent workloads, optional credit packs, and Claude Enterprise, whose access fee excludes usage. Claude Enterprise begins at twenty self-service seats and fifty seats through sales, which keeps it out of reach for a smaller team.6 GitHub Copilot moved to AI-credit billing, but M365 Copilot remains a flat per-seat add-on.4 The distinction changes the budget conversation.

Seat plans need an idle-seat audit. Metered work needs consumption by team, tied to the output it bought and a named owner who can explain the next bill. One audit protects the license pool. The other keeps a useful workflow from becoming an unexplained invoice.

Prompt caching and batch processing matter to the second audit. A cache read costs one tenth of a normal input token, a 90% reduction on repeated cached input. The saving arrives only after repeated context has amortized the cache-write premium. Batch processing cuts both input and output prices by 50% for work that can wait.7 The cost levers are in the plumbing, not in throttling the people doing the work.

A seat count tells you the flat part of the bill. It leaves out the metered work. Track both: cost per active user for seat plans, and dollars per unit of observable output for agent and API workloads. The Q3 2026 briefing covers the split from the board’s side.

Working AI budgets concentrate where output changes

The figures below need to reflect your role mix.

The defensible per-employee figure. For a knowledge-work organization in 2026, total AI spend per employee per month should land somewhere between $40 and $120, with the average closer to $80. Below $40, you’re almost certainly underfunding the people producing leverage. Above $120, you’re almost certainly buying seats for people who don’t use them. The number itself matters less than the shape underneath it.

The 70/20/10 allocation. Roughly seventy percent of your AI dollars should sit on the tools your power users actually choose to open. In practice in 2026, that’s Claude or ChatGPT for chat, and Claude Code, Codex, or Cursor for engineering. Roughly twenty percent goes to workspace AI for the broad middle of your org, where the value is real but modest: drafting in Word, summaries in Excel, meeting recaps. Roughly ten percent is a discretionary API budget for the engineers and operators who turn personal leverage into infrastructure the rest of the org consumes. Adjust the ratios for your role mix. An engineering-heavy org pushes more into category three. A sales-and-marketing org pushes more into category one.

Plan tiers by org size. Under twenty people, buy Team or Business plans on the tools your power users prefer and skip enterprise procurement entirely. Claude Enterprise begins at twenty self-service seats and fifty through sales, so it is unavailable below that threshold.6 From twenty to five hundred people, add enterprise governance where the controls justify it and the vendor’s minimum permits it, then add an API or agent budget for engineers. Five hundred and up, you’ll need vendor-negotiated commits and an SSO story, but resist the urge to consolidate to a single vendor for procurement convenience. Single-vendor lock-in is exactly the trap the bundled-Copilot org fell into. One caveat that was not true a year ago: personal plans buy the tool without governance. Pro and Max run the same scheduled tasks and routines as the business tiers, with no admin surface at all. Team has the admin toggles, but only Enterprise is eligible for the compliance API that turns an agent’s actions into a log your security team can read. That is fine while the work remains personal. Once other people depend on a job, it belongs on a plan you can see into. The seat upgrade is cheaper than the outage when an unowned automation fails silently.

Seats versus API. A seat is the right unit for someone who opens a chat tab a few times a week. API access is the right unit for the engineer building a script that runs ten thousand times a month, and for the operator whose workflow now has a model in the loop on every record. The heaviest few percent of any role should usually have both. The cost of giving them both is small. The cost of forcing them through the seat interface is invisible and large, because they’ll go around you and you’ll lose visibility entirely.

The shape of a working budget. A defensible AI budget for a 200-person mid-market firm in Q3 2026 looks something like this. Twenty to forty Enterprise or Team seats on the tool your power users have chosen, at sixty to one hundred dollars each. A hundred and fifty M365 Copilot seats at the flat $30 price, with the explicit expectation that half will produce modest gains and the other half will be killed in the next audit.4 A two to four thousand dollar monthly API budget, owned by a named engineer, with usage broken out by project. Total: roughly $14,000 to $22,000 a month, or $70 to $110 per employee. That’s a budget you can defend line by line.

AI spend analysis starts with your own AI bill

The phrase is overloaded. This guide covers what your organization spends on AI. “AI in spend analytics” usually means software that uses AI to analyze all of procurement spend. Useful work, different project. A procurement platform cannot solve an AI bill you cannot explain.

Before you can fix the allocation, you have to see it, and most finance systems do not. AI spend hides in several systems. The source tells you where to look and who can control the cost. It does not tell you whether that cost is direct or indirect procurement spend.

AI direct and indirect spend answer the product question

Direct versus indirect is about what the cost does for the thing you sell, not whether an invoice says “AI.” Would this cost disappear if we stopped delivering this product or service? If yes, it is direct AI spend. If the business would keep paying it to run itself, it is indirect AI spend.

Direct AI spend contributes to the product or service a customer buys. Model inference inside a customer-facing product. An AI service delivered to a client. A production cost you can trace to revenue. Indirect AI spend supports the business without becoming the customer offering: internal copilots, recruiting tools, finance automation, support tooling, and employee productivity software.

The same tool can be either. A model API that powers a feature your customer uses is direct. The same API used to summarize internal bug reports is indirect. A consultancy’s Claude subscription may be direct when it produces the client deliverable and indirect when it writes the firm’s own proposals. The tool name is irrelevant. The work it supports decides the class.

Every item needs both classifications. Give it one procurement class, direct or indirect, then one source and visibility class from the four below. The first tells finance how the cost relates to revenue. The second tells the audit where to find it and who can govern it.

Source and representative costWhere it is foundWho sees or controls it
Standalone subscriptions. ChatGPT Business for finance (indirect) or Claude used on client work (direct).AP and GL, SaaS contract register, renewal calendar, SSO seat report.Procurement owns the contract; the business owner and SSO admin can see seats.
Metered consumption. Model inference in a product (direct) or an internal automation (indirect).Cloud and model-provider invoices, project tags, vendor usage dashboard.Engineering controls the workload; finance sees it only when tags and invoices reconcile.
Embedded AI. M365 Copilot for staff (indirect) or AI inside the customer service you sell (direct).Suite renewal, SaaS contract, add-on invoice, product-admin usage report.The suite owner can change seats; procurement sees the renewal; users often see neither cost nor usage.
Shadow or unmanaged AI. A personal account for employee work (indirect) or an unowned client-report automation (direct).Corporate cards, employee expenses, expense reimbursements, automation inventory, and the account itself.The employee or builder controls it until finance, security, or IT finds it.

The analysis is one table of your own. List every source, monthly dollars, procurement class, accountable owner, actual users, and whether usage is visible. Then classify each row as leverage or shelfware: is this spend sitting where the output is, or is it a flat fee on people who never open the tool? A $3,000 model bill can be direct and productive. A $600 collection of personal accounts can be indirect and risky. A $4,500 Copilot deployment can be indirect and mostly shelfware. The source labels do not decide the answer. The work and the usage do.

That classification is the analysis. The audit below is how you pull the raw numbers to fill it in.

Shadow AI exposes weak approved tools

Your security team treats shadow AI as a policy problem. It is also the clearest signal about the quality of your approved tooling.

When your senior salesperson is paying twenty dollars a month out of pocket for ChatGPT Plus while a Copilot seat sits idle in their Microsoft license, they’re telling you something specific. The approved tool is worse than the free tier of the unapproved one. They’ve chosen, with their own money, to route around you.

Run a query against your expense reports for “ChatGPT,” “Claude,” “Anthropic,” “OpenAI,” “Cursor,” and “Perplexity.” Whatever you find is your shadow AI tax, paid in dollars you aren’t capturing and in data exposure you aren’t governing. It’s also the highest-quality user research in your org. Those people did the evaluation for you. They picked the tool. They’re using it to do the work you pay them for.

There’s a newer shape of shadow AI the expense report won’t catch, and it’s the one to worry about now. Your people aren’t only expensing chat. They’re building standing automations: a scheduled task that assembles the Monday report, a routine that reconciles a ledger overnight, an agent running on a clock against systems IT never provisioned it to touch. A pasted prompt was a single risky moment. A scheduled job is durable. It runs after the person who built it changes teams or leaves, and when it stops, or doesn’t stop and nobody remembers why it exists, it is work no one owns. The Q3 2026 briefing calls this shadow AI growing a clock, and it’s the budget signal that now carries the most risk, because the cheap personal plans where these jobs live are exactly the ones with no admin visibility.

The approved tool should be the one people chose, on a plan you can govern, before legal writes the forty-page framework no one will read. Automations that other people depend on need that move sooner, before one fails silently.

A five-minute audit gives you a first pass

This one-page artifact is enough for the next budget meeting.

Five minutes gets you a first pass. A defensible number needs a broader source pull: AP and GL detail, SaaS contracts and renewals, cloud and model invoices, corporate cards and employee expenses, SSO or seat-usage reports, and the scheduled automations people have left running. Do not let employee expenses stand in for the inventory. They reveal only the spend employees were willing or able to submit.

  1. Build the source list. Ask finance for AP, GL, corporate-card, and reimbursement detail. Pull the SaaS contract register and renewal calendar. Then add cloud and model-provider invoices, SSO seat reports, and the automation inventory. Mark each source as standalone, metered, embedded, or shadow or unmanaged.
  2. Pull the seat-level usage report for every AI tool with more than ninety days of seats deployed. Pull it yourself. Do not let the vendor pull it for you. Vendor-supplied dashboards average over the seat pool by design.
  3. Sort by activity. Three buckets. Top ten percent, middle sixty percent, bottom thirty percent. The cutoffs do not need to be precise. The shape will be obvious within sixty seconds.
  4. Pull consumption, not just logins. For metered work, a busy seat is not automatically an efficient one. Pull token or dollar consumption by user and by team from the vendor console and sort it the same way. Where the plan allows it, set per-team and per-user limits, stand up the consumption dashboard, and switch on the per-session cost readout so power users can see what a workflow costs while they run it. The heavy-consumption seats matter as much as the idle ones: a power user is leverage, a runaway script or an uncached workflow is just a bigger invoice.
  5. Find the unmanaged work. Search expenses and corporate cards for the major AI vendors. Ask the teams that own shared systems which scheduled tasks, agents, or scripts call a model. Name an owner for each one. Whatever turns up belongs in the total, even before you decide whether to approve it.
  6. Classify and compute. Give every row a direct or indirect procurement class, then calculate cost per active seat rather than per licensed seat. Take total spend on the tool, divide by seats with meaningful weekly use. Check total AI spend per employee against the $40 to $120 benchmark. Note where the spend is concentrated and where it is sprayed thin.

The output is four numbers, two short lists, and one source inventory. Total monthly AI spend. Cost per active seat. Consumption by team. Shadow or unmanaged AI spend. One list is the bottom thirty percent of seats by activity, the money you free up. The other is the heaviest-consuming seats, the ones you instrument and tune rather than cut. The inventory makes the next quarterly run fast instead of forensic.

The whole exercise takes between twenty minutes and an afternoon, depending on how clean your usage data is. The first time you run it, it will be uncomfortable. The second quarter, it will be a routine.

Your CFO needs a short report and a recurring audit

Three sentences. Edit lightly.

Our total AI spend is roughly $X per employee per month, in line with mid-market benchmarks. We separate direct AI cost in customer delivery from indirect AI cost that runs the business, then concentrate roughly seventy percent of all spend on the tools our highest-output people use every day. Each quarter we audit sources and seat-level usage, kill the bottom thirty percent of inactive seats, and redirect the savings to the people producing measurable leverage.

That paragraph puts a defensible number on the table, shows you understand the difference between spend and allocation, and sets the expectation of a recurring audit. That audit is the governance habit a finance team needs from a budget owner.

If your CFO pushes for a traditional ROI model, explain that the instrument does not measure what AI actually does. AI spend produces cycle-time compression, capability expansion, and the elimination of work that previously sat in a backlog. None of that lands cleanly on a P&L line designed for cost takeout. A quarterly review of the audit numbers, plus three named workflows that have measurably changed since the last review, is a better scoreboard.

Four purchasing choices will waste money in Q3 2026

A few specific things that will cost you money in Q3 2026 if you don’t watch for them.

A multi-year usage commitment needs a full quarter of use first. Metered work can grow rapidly once a workflow lands. If you commit to $50,000 a month and use $20,000, you’re paying the difference. Commits should follow the ramp, not lead it.

M365 Copilot is not already paid for. It is a thirty-dollar per-seat add-on, which now puts list price around $69 on E3 and $90 on E5.4 The sunk-cost framing is a vendor tactic. The all-in cost belongs in the seat audit.

Enterprise agent platforms still need pilots. All four major vendors shipped one in April. None of them are mature. A production workflow needs to run to completion before a platform purchase. A platform selected under Q3 pressure will be hard to defend in Q4.

Single-vendor consolidation remains a procurement convenience. The market is stratifying. Know which tier each vendor is selling you, and keep at least two of them honest with each other.

AI spend terms need a common meaning

AI spend analysis gives you one view of AI costs

It is the work of finding and analyzing what your organization spends on AI. This is different from “AI in spend analytics,” which means using AI to analyze all procurement spend. An AI spend analysis pulls subscriptions, consumption, embedded features, and unmanaged tools into one view, then tests each line for direct or indirect classification, visibility, and leverage.

Direct and indirect spend depend on the work, not the tool

Direct AI spend contributes to the product or service you sell: inference in a customer-facing product, an AI service delivered to a client, or a production cost traceable to revenue. Indirect AI spend supports the business: internal copilots, recruiting, finance automation, support tooling, and employee productivity software. Ask: Would this cost disappear if we stopped delivering this product or service? The answer determines the class.

AI spend hides across internal sources

Start with AP and GL detail, SaaS contracts and renewals, cloud and model invoices, corporate cards, employee expenses, SSO or seat-usage reports, and scheduled automations. Those are visibility classes, not procurement classes. Each row still needs the direct-or-indirect test and an accountable owner.

Spend management is a recurring allocation discipline

Spend management is the recurring discipline of watching that allocation. It means running the seat-and-consumption audit each quarter, killing the bottom slice of idle seats, setting per-team and per-user spend limits on the metered tools, and redirecting the freed dollars to the people producing measurable output. The single most important habit a finance team can hear from a budget owner is that this happens every quarter.

The audit combines source records, use, and consumption

Run the five-minute audit below. Pull the source records and seat-level usage report yourself, sort activity into power, occasional, and idle buckets, pull consumption by user and team, compute cost per active seat rather than per licensed seat, find shadow and unmanaged work, and check the per-employee total against the $40 to $120 benchmark. The output is four numbers, two short lists, and a source inventory that fits on one page.

One seat audit reveals the next reallocation

Pull the seat-level usage report for whichever AI tool you have the most seats of. If that’s Microsoft 365 Copilot, you’re about to have an uncomfortable hour. Sort by last-thirty-day activity. Identify the bottom thirty percent of seats. Send a note to those users. Two weeks to demonstrate use, or the seat goes back into the pool.

Take the dollars you free up and offer them, in the same week, to the three people in your org you suspect are getting the most leverage from AI right now. Ask what tool they would actually choose. Buy that tool. Track what changes.

That audit and reallocation show more about the shape of your AI spend than the next vendor pitch deck. They also give your CFO a defensible account for the quarter.

Footnotes

  1. OpenAI’s ChatGPT Business pricing lists $20 per user per month on annual billing and $25 month to month. OpenAI says the revised structure took effect April 2, 2026. Last verified August 4, 2026.

  2. Microsoft reported more than 20 million paid M365 Copilot seats in Q3 FY26. Microsoft put its paid M365 commercial base above 450 million in its FY26 Q2 earnings materials. Last verified July 29, 2026.

  3. Joe Salesky, “AI Choice 2026: Why Licenses Don’t Equal Adoption”, Recon Analytics, February 3, 2026. The U.S. paid-subscriber panel published readings for July 2025, September 2025, and January 2026; it did not publish an April reading. Last verified July 29, 2026.

  4. Microsoft says its commercial M365 suite changes took effect July 1, 2026 in its pricing and packaging update. M365 Copilot remained a $30 add-on, making annual-commitment list price $69 on E3 and $90 on E5. Microsoft listed Copilot Business at $18 promotional and $21 list for organizations with up to 300 seats through September 30, 2026. Last verified July 29, 2026. 2 3 4

  5. Anthropic lists Claude Team at $20 per seat monthly on annual billing and $25 monthly. OpenAI lists ChatGPT Business at $20 per user monthly on annual billing and $25 monthly. Microsoft lists M365 Copilot as a $30 per-user add-on, and Google lists Workspace Standard at $14 per user monthly. These business and workspace plans include a defined allowance in their seat price. Last verified July 29, 2026.

  6. Anthropic’s pricing page lists Claude Enterprise as access plus API-rate consumption, without included usage. Its self-service path starts at 20 seats; sales-assisted Enterprise starts at 50 seats. Last verified July 29, 2026. 2

  7. Anthropic’s API pricing lists cache reads at 0.1x base input tokens, five-minute cache writes at 1.25x, one-hour writes at 2x, and Batch API input and output at 50% off. Last verified July 29, 2026.