TL;DR. AI tool value is concentrated in the few percent of your people who extract it. Procurement-led selection optimizes for the median user, far from the leverage. Split the decision. Standardize the durable foundation, the API, agent platform, and internal tooling you expect to outlast this year’s model, on Anthropic. Let the chat layer stratify: leveraged users pick the tab they open, and you fund what they already use. Standardize the rest after the market inside your walls converges.
You have an AI tool. Probably more than one. There is a renewal coming up. There is a vendor on your calendar. There is an internal champion arguing for the platform you don’t have. There is a peer at another company who told you, in a tone that implied you were behind, that they “standardized on” something.
None of those conversations are improved by another vendor demo. The decision you actually have to make is older and simpler than the marketing makes it sound. Which tool, for which roles, at which tier. The reason it feels hard is that almost every method commonly used to make this decision is wrong.
Procurement selects for the median user
The default playbook for a SaaS purchase. Run an RFP. Build a feature matrix. Negotiate with the vendor with the best partnership program. Standardize on one. Push it to everyone. Train everyone. Measure adoption. Move on.
Payroll software has to work the same way for every employee. AI tools are used hard by a small fraction of your workforce and ignored by most of the rest, while the productivity gain appears in the work the leveraged few can now do. Procurement logic optimizes for the wrong end of that distribution.
You can already see this in your own org. Your most senior leveraged user is paying for ChatGPT Plus on a personal card and pasting work into it because the approved tool is worse than the free one. Your engineers built their own internal Claude Code workflow because waiting for the procurement decision was going to cost them a quarter. Your marketing director runs a personal Claude account on the side and uses it for the work she actually wants to ship. The tools your org bought are not the tools your org’s leverage runs on.
It’s a selection problem. The procurement process selected for the wrong qualities, in the wrong vendor, against the wrong success metric.
Six procurement habits pick the wrong tools
Six variations recur.
“We already pay for Microsoft, so Copilot.” The most common one. M365 Copilot has more than 20 million paid seats, which sounds large until you notice it is roughly 4.4 percent of the 450 million-plus M365 commercial base.1 The preference data matters more than that ratio. When workers can choose between ChatGPT, Gemini, and Copilot, 70 percent choose ChatGPT, 18 percent Gemini, and 8 percent Copilot. When Copilot is the only option, 68 percent adopt it.2 That is a product-preference problem, not a deployment problem. Among lapsed users, 44.2 percent cite distrust of Copilot’s answers, compared with 42.8 percent for Gemini and 40.6 percent for ChatGPT.3 The gap is only a few points. It is still the wrong chat tool to make your only provision.
“We ran an RFP.” RFPs select for the vendor that maps best to a fixed feature list. AI tool value is in qualities that do not fit RFP rows. How the model reasons. How it handles being interrupted and corrected. What its refusal behavior looks like. How it feels to keep open all day for a month. None of these fit a checkbox. The RFP finalist is almost never the tool your power users will actually open.
“Let’s standardize on one vendor for everything.” Standardizing the whole stack on one name is a late-stage move, and no single all-purpose winner has emerged. But the picture has sharpened since the start of the year. The durable foundation has consolidated: for anything you mean to build and keep, the API, the agent platform, internal tooling, custom workflows, Anthropic is now the enterprise default. The chat layer is what still stratifies. OpenAI leads chat preference. Microsoft leads embedded gains in spreadsheets and documents. Google leads Workspace-native shops with infrastructure ties. So standardize the foundation, and let the chat tab be a preference. Forcing one vendor across both means accepting the worst experience in most of the categories your people work in. The full picture is in The State of AI: Q3 2026, and the head-to-head buyer’s read on the three names people actually compare is ChatGPT vs. Claude vs. Gemini.
“Let’s pick an agent platform.” Every major vendor shipped an enterprise agent platform on or around April 22.4 All of them want to be the layer your company builds custom agents on for the next decade. None of them are mature. Reliability evidence is product- and workflow-specific, not a market-wide score. A platform choice now locks your company into one vendor’s current failure modes. A vendor urging platform standardization in 2026 is asking you to adopt theirs.
“We hired a consultant to choose.” The consultant runs the same RFP.
“We’re letting it grow organically.” Organically usually means your people are pasting client data into free ChatGPT because the approved tool is worse than the free one. Shadow AI is a top-down provisioning failure. Better approved tools stop people from working around them.
The thing all six have in common is that they treat AI tool selection as a procurement exercise to be done by the people who do procurement, on behalf of the people who use the tools. This is exactly the inversion that produces the wrong outcome.
Leverage sits with the users procurement ignores
Here is the underlying picture, in one paragraph.
Your AI productivity is lopsided. A small group of users, call it the top five percent by consumption, is running ten times the volume of the median and a meaningful multiple of the output. The rest of your seat pool produces something between zero and a small marginal gain. The aggregate productivity number is the average of those two populations. A procurement decision made on the aggregate is, in effect, a decision made on behalf of the larger population, which is the population producing the smaller fraction of the value. The smaller population, the one producing most of the value, has very specific tool preferences. Overriding those preferences downgrades the only people whose output your AI program is actually moving.
Recognizing Leverage covers the disposition that produces this distribution and the signals you can use to find the people in it. Once you know where the value sits, purchase around it.
Tool access should follow three tiers
The framework has three rings, deliberately differentiated by who is in them and what they get.
Tier 1 gives everyone a strong chat tool
Everyone in the org should have a baseline chat tool: the one your most leveraged users actually open, not the one already in the contract.
In most orgs in Q3 2026, that is ChatGPT Business or Claude Team. Both cost $20 per user per month on annual billing and $25 month to month.5 Claude Team also has a Premium seat at $100 per seat per month on annual billing or $125 monthly for the people who use Claude Code or Cowork hard. ChatGPT Go at $8 per person per month is a low-cost option for a small group before an Enterprise deal; use it where central administration is not yet the requirement. Both baseline tools are better daily-driver chat options than M365 Copilot. Pick one. If you cannot pick one, give people both. The cost of running two baseline chat subscriptions is rounding error against the cost of the wrong one.
Tier 1 is for drafting, brainstorming, analysis, and the daily knowledge-work tasks where chat is the interface. Embedded document gains and coding belong in different tiers.
Most people in Tier 1 will never become heavy users. A good baseline gets the broad middle out of free ChatGPT and gives the dispositionally inclined, before anyone has spotted them, a real tool when they start to use it.
Tier 2 gives roles the tools that fit their work
Beyond the baseline, specific roles get specific tools, chosen by the people doing the role, not by procurement.
Engineering gets a coding agent. Claude Code, Codex, Cursor, GitHub Copilot. The choice belongs to the engineers, not to the CIO. The tools are stratifying by preference and the right answer is to fund whichever the team converges on, often more than one. Coding agent budget is the highest-return line item in an AI program. Do not ration it. Menlo Ventures put Anthropic at 54 percent of enterprise coding spend in its December 2025 survey.6
Doc-heavy roles (finance modeling in Excel, contract drafting in Word, ops in spreadsheets) get M365 Copilot, or Gemini in Workspace, depending on which suite they live in. Test the embedded workflow in the work itself. This is what Copilot is actually for. The preference data says it needs a Tier 1 chat tool beside it.2
Customer support gets the support-specific AI built into their platform, not a chat tool retrofitted into a help desk. Zendesk AI, Intercom Fin, Salesforce Agentforce as it matures. The integration with ticket history and knowledge base is the value. A bare chat tool does not have either.
Sales gets call-prep and account-research workflows built on top of Tier 1, plus whatever their CRM ships. Most “sales AI” is repackaged chat with a vector DB of LinkedIn profiles attached. Tier 1 plus a thirty minute prompt-template session does most of the same job for none of the markup.
Legal gets the baseline plus a contract-review specialist (Harvey, Spellbook, others) once the volume justifies the cost. Below a certain document throughput, the specialist is overpriced for the use; above it, the specialist pays for itself in associate hours saved.
The pattern across roles is the same. The embedded or specialty tool wins where the workflow is fixed and the integration is the value. The chat tool wins where the work is open-ended.
Tier 3 funds power users without rationing
The few percent producing the leverage should have whatever they ask for.
Multiple chat subscriptions, because Claude and ChatGPT are good at different things and a power user who knows both will use both. Direct API budget for whatever they are building. ChatGPT Pro at $200 a month for the people who use its full allowance, or a separate Codex seat where credit-based usage earns it.7 The specialty tool they identified at the last bake-off. A second seat of something for the prototype they want to run in parallel.
The calculation is simple. The output of a leveraged user dwarfs their annual tool budget. If you are negotiating with them about whether they can have Cursor and Claude Code, leadership attention is in the wrong place.
This tier also shows which tool is winning inside your walls, the input to standardization later. Power users gravitate toward what works. Their choices supply the evidence.
Each serious vendor belongs in a different tier
The market has stratified enough that the question is which vendor belongs where.
Anthropic (Claude). The enterprise default, and the place to standardize anything you mean to keep. Menlo Ventures put Anthropic at 40 percent of enterprise LLM API spend and 54 percent of enterprise coding spend in its December 2025 survey.6 Use it as the foundation: the API, the agent platform, the internal tooling, the custom workflows you want to outlast this year’s model. Its Agent Skills format, the packaging that turns a model into something that knows your finance close or your contract review, was opened as a standard this year, and a plug-in ecosystem grew up around it for finance, legal, accounting, and data science. That makes Anthropic a defensible procurement choice. Claude Code is still the proven coding force multiplier. Claude Team remains a flat seat plan with included usage. Enterprise adds access plus consumption and starts at 20 self-service seats, which makes it a governance choice rather than the default for a small team.8 Anthropic is weaker as the chat tab. For image generation, multimodal breadth, and the tool your non-technical people open without being told, look elsewhere.
OpenAI (ChatGPT, Codex). OpenAI belongs in the chat layer: the default tool your people open without prompting, the most comfortable interface for non-technical users, image generation, and multimodal breadth. This is the preference to let your leveraged users exercise. The chat tab is exactly where a house standard does not belong. Codex is a serious enterprise coding agent, not a sidecar: it has workspace roles and seats, local runtime policy for desktop, CLI, and IDE use, cloud environment and repository controls, analytics, and a Compliance API.9 Anthropic, rather than OpenAI, is now the foundation for durable agents and internal tooling. The agent layer is churning. Workspace Agents is a research preview for Business, Enterprise, and Edu; it moved to credit-based billing in July, and Agent Builder shuts down November 30.10 A low-stakes async pilot produces enough information without a platform commitment.
Microsoft (Copilot). Microsoft is worth buying for the embedded features in Excel, Word, and Outlook, where the workflow benefits from suite context. GitHub Copilot can serve engineering teams that have not picked something else. GitHub Copilot moved to AI-credit billing on June 1, 2026. M365 Copilot did not. It remains a flat per-seat add-on.11 Copilot loses the chat preference test when workers have a real choice.2 M365 Copilot belongs beside a chat tool, not in place of one.
Google (Gemini). Google fits Workspace-native shops where everyone already lives in Docs, Sheets, and Gmail, and infrastructure-heavy organizations with Google Cloud as their primary platform. The new Gemini Enterprise Agent Platform supports Anthropic’s Claude models alongside Gemini, which is a useful signal about where Google thinks the market is going. Test it against Claude Code, Codex, and Cursor before using it as a coding tool. The preference test favors ChatGPT for consumer-facing chat.2
The chat layer is stratifying; the foundation has consolidated on Anthropic. Assign each vendor to a tier and standardize only the part ready for it.
M365 Copilot should stay with embedded users
Many organizations have the Microsoft problem.
You bought M365 Copilot for everyone. Most of the seats are idle. The renewal is coming up. Your Microsoft rep is pitching you Copilot Studio and trying to get you to standardize on the agent platform.
The framework, in five steps:
- Pull the seat-level usage report. Identify the bottom 30 percent.
- Of the active users, distinguish the embedded-feature users (Excel, Word, Outlook) from the chat users.
- Keep Copilot for the embedded-feature users. Cancel for everyone else.
- Redirect the saved budget to ChatGPT Business or Claude Team for the chat use case.
- Tell your Microsoft rep your renewal will be smaller this year, and that you would be open to a re-pitch when chat NPS goes positive.
The usage report makes this move easy to defend with a CFO. Your Microsoft rep may not be happy. That tells you whose interest the rep represents. The mechanics of the audit itself (pulling the seat report, what to look at, how to defend the cuts) are in Evaluating Spend.
The conversation gets harder when someone at the executive level made the original Copilot decision. Frame it around what new usage data changes: “the data we did not have at the time of the decision is now available, and the cost of not acting on it is X dollars per quarter.” That is defensible at any altitude.
Agent platforms need pilots, not commitments
Every vendor wants you to commit to their agent platform. OpenAI’s Workspace Agents. Google’s Gemini Enterprise Agent Platform. Microsoft’s expanded Copilot Studio. Salesforce Agentforce in partnership with Google. The April 22 launches were a coordinated land grab.
The Q3 2026 move is to pilot without committing.
Three reasons. First, reliability is specific to the product and workflow. A test of one agent does not give you a market-wide failure rate. Your own test has to answer which workflow fails, how it fails, and who catches it. Second, the platforms are moving too quickly for a durable commitment. Third, the transferable asset is the ability to specify agent workflows: what you want the agent to do, what the success criteria are, and what the escalation path looks like when it fails. That ability transfers across platforms. The platforms will not.
Use two or three low-stakes workflows: status reports, expense categorization, and internal data lookups. Run each on a different vendor’s agent platform, measure failure rates, keep notes, and re-evaluate quarterly.
Scheduled tasks are one capability ready for real use. Cowork introduced them on February 25, reached desktop general availability on April 9, and arrived on web and mobile on July 7.12 A prompt and a cadence are enough. They suit recurring work with a checkable output: a Monday status brief assembled from five systems, a weekly competitive scan, or a daily reconciliation that flags what did not tie out. The person who used to spend Monday morning compiling the report reads a draft waiting when they sit down.
The operating constraint changed. Cloud sessions are beta, but they continue with the laptop closed and scheduled tasks can run with no device online. Work that needs local files or local apps still needs an awake machine. Cowork conversation history on local machines cannot be centrally managed or exported by admins, even on Enterprise. Cloud Cowork is the route that reaches the Compliance API. A scheduled job can now keep running when no one has a laptop open, making ownership, access, logs, checks, and shutdown more urgent. Every recurring task needs a named owner and a clear location: local or cloud.12
The verifier constraint still governs. Schedule only work that two people can mechanically agree is right. Work that needs a careful senior read before anyone trusts it is not ready for a cadence. Full platforms remain for pilots; scheduled tasks are ready for use.
If your CIO is being told to “pick a platform” by Q3, the answer is no. The platform race is open, and the cost of being wrong is higher than the cost of waiting.
Standardize after power users converge
The timing is the problem. Premature standardization is bad.
One layer is ready. The foundation for durable agents and internal tooling has converged on Anthropic, and standardizing it now is defensible. The chat and role tools still need to converge inside your walls.
The signal that says you are ready to standardize: power users from different starting tools have converged on the same one without being told to. Three engineers picked Cursor on their own. Two product managers are quoting Claude in Slack. The marketing team’s leveraged user is forwarding ChatGPT outputs to her director. When two or three of your highest-leverage users in a role are independently picking the same tool, that tool has won inside your walls. Then standardize. Negotiate the enterprise contract. Roll it out broadly. Train the broad middle on a tool that has already proved itself.
The reverse, picking the tool first and trying to make it the one everyone converges on, is the procurement default. It does not work in this market. The market is moving too fast and the user preferences are too strong.
Until you see convergence, running two or three tools is information gathering at low cost. The information you are gathering is which tool wins inside your walls, which is the only basis for a defensible standardization decision later.
Five rules keep tool spend tied to leverage
Five rules a director can screenshot.
- Buy the tool your leveraged users already open. Not the one already in your contract.
- Standardize the foundation now; stratify the chat layer. Build durable agents and internal tooling on Anthropic. Run two or three chat tools until your power users converge on one.
- Do not ration power users. Their tool budget is rounding error against their output.
- Cut idle seats every quarter. Redirect the savings to tools people actually use.
- Pilot agent platforms, do not pick one. The platforms are not mature. The workflow muscle is what carries.
Tool spend that follows these five rules gives the rest of the program room to work. A renewal gives you the moment to correct it.
A renewal starts with your power users
Before your next AI vendor renewal, ask the three most leveraged AI users you can name which two tools they would pick if it were their decision. Not “which tool do you like.” Which two would you pick. The framing matters. You are looking for the tools that survive an honest comparison, not the tools that have a small advantage in one demo.
Those answers determine the first reallocation: fund the two tools they pick and cancel seats for the one they leave behind. Power-user choices supply the data a renewal needs.
If you can’t name three leveraged users, that’s a different problem. Recognizing Leverage is where to start.
Footnotes
-
Microsoft reported more than 20 million paid M365 Copilot seats in Q3 FY26, up more than 250 percent year over year. Its Q2 FY26 disclosure put the commercial M365 base above 450 million seats. Last verified August 14, 2026. ↩
-
Joe Salesky, Recon Analytics, “AI Choice 2026: Why Licenses Don’t Equal Adoption”, February 3, 2026. Among workers with all three tools, 70 percent chose ChatGPT, 18 percent Gemini, and 8 percent Copilot; when Copilot was the only employer-provided option, 68 percent adopted it. The report draws on 150,000-plus U.S. paid AI-subscriber respondents. Last verified August 14, 2026. ↩ ↩2 ↩3 ↩4
-
Recon Analytics, “AI Choice 2026: Why Licenses Don’t Equal Adoption”, February 3, 2026. Among users who tried and stopped using each tool, 44.2 percent cited distrust of Copilot’s answers, against 42.8 percent for Gemini and 40.6 percent for ChatGPT. Last verified August 14, 2026. ↩
-
OpenAI released Workspace Agents, Google announced Gemini Enterprise Agent Platform, and Salesforce and Google announced their Agentforce integration on April 22, 2026. Google’s Model Garden carries Claude models alongside Gemini. Last verified August 14, 2026. ↩
-
Anthropic lists Claude Team Standard at $20 per seat monthly on annual billing and $25 monthly, and Premium at $100 per seat monthly on annual billing or $125 monthly. OpenAI lists ChatGPT Business at $20 per user monthly on annual billing and $25 monthly. ChatGPT Go costs $8 per month in the U.S.. Last verified August 14, 2026. ↩
-
OpenAI lists ChatGPT Pro at $200 per month. Its Codex rate card covers separate credit-based Codex usage and says costs vary by workload. Last verified August 14, 2026. ↩
-
Anthropic’s pricing page lists Claude Team for 2 to 150 people and Claude Enterprise at $20 per seat plus usage at API rates. Enterprise starts at 20 self-service seats and 50 through sales. Last verified August 14, 2026. ↩
-
Codex became generally available on October 6, 2025. OpenAI’s July 9, 2026 release notes placed Chat, Work, and Codex in one desktop app. Its enterprise admin guide documents workspace roles, managed configuration, cloud controls, analytics, and the Compliance API. Last verified August 14, 2026. ↩
-
OpenAI’s Business and Enterprise rate card says Workspace Agent runs are credit-based and charged by token usage. Agent Builder’s documentation says it will shut down November 30, 2026. Last verified August 14, 2026. ↩
-
Microsoft’s pricing and packaging update keeps M365 Copilot at $30 per user monthly on an annual commitment. GitHub documents the separate AI-credit model for Copilot premium requests. Last verified August 14, 2026. ↩
-
Anthropic’s release notes record scheduled tasks on February 25, 2026, Cowork desktop general availability on April 9, and web/mobile Cowork on July 7, when remote sessions began beta rollout and could keep working with no device online. The Cowork enterprise guide says local conversation history cannot be centrally managed or exported by admins. Last verified August 14, 2026. ↩ ↩2