# Forecasting and Controlling the Cost of Enterprise AI

### A field guide for CEOs, CHROs, and CFOs — with the technical depth CTOs, CISOs, and CIOs need

*Prepared for the Argentum AI Leadership Sprint. Pricing verified against provider pages and analyst sources dated 2025–2026. All figures are USD list prices unless noted; enterprise contracts are negotiated and will differ.*

---

## The one idea to hold onto

For thirty years, enterprise software cost was a **seat problem**: you counted employees, multiplied by a price, and signed a three-year contract. Your bill was flat and predictable, whether a user logged in once a month or lived in the product all day.

AI is quietly ending that arrangement. The seat isn't going away, but a **second meter** is being bolted on underneath it — one that charges for *work done*, not *people licensed*. Microsoft calls the new unit a "Copilot Credit." OpenAI calls it a "credit." Anthropic calls it a "token." They are all the same idea: **you now pay for consumption, and consumption is driven by machines, not headcount.**

This matters because the two meters behave in opposite ways. Seat cost is **linear and capped** — 1,000 people cost twice as much as 500. Consumption cost is **non-linear and, by default, uncapped** — a single misconfigured agent running overnight can spend more in eight hours than a department of humans spends in a month. The MIT Media Lab's widely-cited 2025 study found that **95% of enterprise generative-AI pilots had produced no measurable financial return** ([Fortune](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/)), and one field review of 120 deployments found **27% exceeded their budgets, by an average of 38%** ([Fazen Capital](https://www.fazen.markets/en/ai-cost-cutting-risks-2026)). The cost problem is not that AI is expensive; it is that AI cost is **structurally harder to forecast** than anything on your P&L today.

The rest of this report explains what you can know now, what's coming, how to budget for it, and the specific control levers that keep the second meter from surprising you.

---

## Part 1 — What leaders can understand NOW

### 1.1 The four pricing models, in plain English

There are only four billing mechanics in enterprise AI. Every product you evaluate is a combination of these.

| Model | How you pay | Predictability | Who uses it |
|---|---|---|---|
| **Per-seat subscription** | Fixed $/user/month | High — flat and capped | ChatGPT Business, M365 Copilot, Claude Team |
| **Token / usage pricing** | $ per million words in + out | Low — scales with activity | All provider APIs (OpenAI, Anthropic, Azure) |
| **Credit systems** | Prepaid or metered "credits" per action | Medium — depends on action mix | Copilot Studio, ChatGPT agent features |
| **Subscription + credit hybrid** | Flat seat *plus* metered overage | Low above the included allowance | The 2026 direction for all three vendors |

**Per-seat** is the model executives already understand. You buy a license, the person uses it as much as they want (within fair-use limits), the bill is flat. This is still how the flagship assistants — ChatGPT Business, Microsoft 365 Copilot — are sold to *humans*.

**Token pricing** is what powers everything underneath. A "token" is roughly ¾ of a word. Providers charge separately for **input** (what you send the model — your prompt plus any documents and conversation history) and **output** (what it writes back). Output is typically **4–5× more expensive** than input. Anthropic's own convention is that output runs a fixed 5× its input rate, so a useful mental shortcut is: *know your input cost, multiply by five for output* ([BenchLM](https://benchlm.ai/blog/posts/claude-api-pricing)).

**Credit systems** sit in between. Instead of exposing raw tokens, the vendor invents a synthetic unit — a "credit" — and charges a different number of credits per *type* of action. A simple lookup costs 1 credit; a data-grounded answer costs 10; an autonomous agent action costs 25+. This lets vendors price *behavior* rather than *compute*, and it's where the industry is heading for agents.

**Hybrids** combine a flat seat (which includes an allowance) with metered charges once you exceed it. This is the model to watch: it *looks* like predictable per-seat pricing but *behaves* like usage pricing at the margin.

### 1.2 What actually drives cost — the five levers

Every dollar of AI cost is some combination of these five drivers. Non-technical leaders should be able to name all five, because every control strategy in Part 4 targets one of them.

1. **Seats** — how many people are licensed. Linear, predictable, the easy part.
2. **Tokens** — the total volume of text flowing in and out. This is the real fuel gauge.
3. **Model choice** — which "brain" you invoke. The gap between a cheap and a premium model is enormous: on Anthropic's rate card, the flagship Opus tier costs **$5 input / $25 output per million tokens** while the small Haiku model costs **$1 / $5** — a **5× difference for the same task** ([Claude Platform Docs](https://platform.claude.com/docs/en/about-claude/pricing)). OpenAI's spread is even wider: GPT-5.5 at **$5 / $30** vs. GPT-5.4-nano at **$0.20 / $1.25** — roughly **25×** ([OpenAI API pricing](https://developers.openai.com/api/docs/pricing)).
4. **Context length** — how much history and reference material you stuff into each request. Because you pay for input tokens *every single call*, a long system prompt or a large attached document is a tax you pay repeatedly. (The good news: the newest Anthropic and OpenAI flagships now include a **1-million-token context window at standard pricing**, ending the old premium surcharge for large contexts — [Claude Platform Docs](https://platform.claude.com/docs/en/about-claude/pricing).)
5. **Agent runtime** — how long an autonomous process runs and how many model calls its loop makes. This is the driver that scales non-linearly and is covered in depth in Part 4.

The critical insight for a CFO: **seats are what you sign for, but tokens are what you pay for.** A per-seat contract hides drivers 2–5 inside a flat number. The moment you move to agents, those drivers surface as a separate, variable line item.

### 1.3 The vendor pricing landscape today (2025–2026)

**OpenAI — ChatGPT for organizations**

| Plan | Price | Notes |
|---|---|---|
| **Business** (formerly Team) | **$20/user/mo annual, $25 monthly** | 2-seat minimum; price cut $5 on Apr 2, 2026 ([OpenAI Help](https://help.openai.com/en/articles/8792536-managing-billing-and-seats-in-chatgpt-business)) |
| **Enterprise** | **~$60/user/mo** (unpublished; $45–$75 range) | 150-seat minimum, annual prepaid → ~$108K/yr floor ([Beam Cloud](https://www.beam.cloud/blog/chatgpt-enterprise-pricing)) |

**Microsoft 365 Copilot** — the licensing maze (see 1.4)

| SKU | Add-on price | Seat cap |
|---|---|---|
| **Copilot Business** | **$21/user/mo** annual (was promo $18) | ≤300 users |
| **Microsoft 365 Copilot** (Enterprise) | **$30/user/mo** annual | No cap |

**Anthropic — Claude for organizations**

| Plan | Price | Notes |
|---|---|---|
| **Team Standard** | **$20/seat/mo** annual ($25 monthly) | 5-seat minimum ([Claude pricing](https://claude.com/pricing)) |
| **Team Premium** | **$100/seat/mo** annual | Includes Claude Code developer tools |
| **Enterprise** | **Seat price + usage at API rates** | Admin spend limits, SSO, SCIM, audit logs ([Claude pricing](https://claude.com/pricing)) |

Note Anthropic's Enterprise structure explicitly reads *"Seat price + usage at API rates"* — the hybrid model, stated on the pricing page. In April 2026 Anthropic **stripped bundled tokens out of its enterprise seat deal**, moving renewing customers onto usage-based plans ([The Register](https://www.theregister.com/software/2026/04/16/anthropic-ejects-bundled-tokens-from-enterprise-seat-deal/5226555)). This is the clearest signal yet of where the industry is going: the flat seat is being unbundled from the consumption it enables.

### 1.4 The Microsoft Copilot licensing maze (a special case)

Microsoft Copilot is where most enterprises will meet AI cost complexity first, because the sticker price is deeply misleading. **The $30 add-on is never your real cost**, for three reasons.

**Reason 1 — Copilot is an add-on, not a product.** It requires a qualifying base Microsoft 365 license underneath it. Your true per-user cost is *base + add-on*:

| Path | Base license | + Copilot | **True cost/user/mo** |
|---|---|---|---|
| SMB (Business Standard) | $12.50 | $21 | **$33.50** |
| SMB (Business Premium) | $22 | $21 | **$43** |
| Enterprise (E3) | $36 | $30 | **$66** |
| Enterprise (E5) | $57 | $30 | **$87** |

*Source: [GoSearch](https://www.gosearch.ai/blog/microsoft-copilot-pricing/), [EPC Group](https://www.epcgroup.net/blog/microsoft-365-copilot-pricing-licensing-enterprise-guide-2026).*

For a 10,000-person enterprise on E3, the all-in Copilot bill is roughly **$8.28 million a year**, not the $3.6M the $30 figure implies ([LinkedIn analysis](https://www.linkedin.com/pulse/microsoft-copilot-m365-what-30-per-user-price-actually-iqe1f)).

**Reason 2 — the base prices are rising.** Microsoft announced in December 2025 that suite prices increase July 1, 2026: **E3 rises $36→$39, E5 $57→$60** ([SAMexpert](https://samexpert.com/microsoft-365-copilot-licensing/)). An E3 enterprise adding Copilot will pay **$69/user/mo** from mid-2026 — for 5,000 users, **$4.14M/year** ([LinkedIn](https://www.linkedin.com/pulse/microsoft-copilot-m365-what-30-per-user-price-actually-iqe1f)).

**Reason 3 — agents are metered separately and are NOT in the seat price.** This is the trap. The Copilot *seat* covers a human using Copilot in Word, Excel, Teams. But the moment you build a custom **agent** in Copilot Studio, you enter a completely different, consumption-based billing world (covered in Part 2). A licensed Copilot user gets certain agent actions **zero-rated** (free), but anything beyond that draws down metered credits ([A Guide to Cloud](https://www.aguidetocloud.com/blog/copilot-credits-explained/)).

**The practical takeaway for procurement:** on the small-business lane, the *bundled* "with Copilot" SKUs are often cheaper than buying the pieces — Business Standard + Copilot is **$23.50** bundled vs. **$30.50** unbundled ([Velosio](https://www.velosio.com/blog/m365-copilot-pricing-calculator/)). And a "Teams-free" variant of the base plans shaves a few dollars per seat. The maze rewards attention.

---

## Part 2 — What to EXPECT next from Microsoft, OpenAI, and Anthropic

The direction is unanimous and unmistakable: **all three vendors are layering consumption-based agent billing underneath their flat seats.** The seat pays for a *person*; a new meter pays for *machine work*. Here is the current state of each meter — with real numbers.

### 2.1 Microsoft: the "Agent Economy," Copilot Credits, and Agent 365

Microsoft has gone furthest in formalizing a machine-work meter. On **September 1, 2025**, it renamed its agent billing unit from "messages" to **Copilot Credits** and moved from a flat per-message tally to **feature-based rates** — different actions cost different amounts ([Microsoft Learn](https://learn.microsoft.com/hr-hr/microsoft-copilot-studio/billing-licensing)).

**A Copilot Credit is worth ~$0.01.** You buy them two ways:
- **Pay-as-you-go:** $0.01/credit, billed through Azure, no commitment ([Microsoft Learn](https://learn.microsoft.com/hr-hr/microsoft-copilot-studio/billing-licensing)).
- **Prepaid capacity pack:** **$200 for 25,000 credits/month** per tenant (≈$0.008/credit) — but **no rollover**; unused credits are lost ([Reveal Compliance](https://revealcompliance.com/blog/copilot-studio-pricing)).

**What each agent action costs** (from Microsoft's own billing docs):

| Agent action | Copilot Credits | ≈ USD |
|---|---|---|
| Classic answer (scripted) | 1 | $0.01 |
| Generative answer (LLM-composed) | 2 | $0.02 |
| Agent action (a step/tool call) | 5 | $0.05 |
| Tenant Graph grounding (uses your data) | 10 | $0.10 |
| Agent flow actions (per 100) | 13 | $0.13 |
| Autonomous agent trigger | ~25 | ~$0.25 |
| Premium AI tools (deep reasoning) per 10 responses | 100 | $1.00 |

*Source: [Microsoft Learn billing rates](https://learn.microsoft.com/en-us/microsoft-copilot-studio/requirements-messages-management), [SAMexpert](https://samexpert.com/microsoft-365-copilot-licensing/).*

The crucial detail: **a single user message can cost anywhere from 1 to 200+ credits** depending on how the agent is architected ([Reveal Compliance](https://revealcompliance.com/blog/copilot-studio-pricing)). A support agent grounded in your data with actions realistically runs **12–22 credits per turn** ([Frontrow](https://frontrowtech.com.au/insights/copilot-studio-message-pack-pricing-australia-2026)). At 6 turns per conversation, that's ~$1.00–$1.32 per conversation — small individually, dangerous at volume.

**The near future: Agent 365 and the E7 "Frontier" tier.** Microsoft has introduced a new top tier, **Microsoft 365 E7 at ~$99/user/mo** (E5 + Copilot + Agent 365 + Entra Suite), which buys **full autonomous multi-step agent execution** rather than the "suggest-and-approve" model of Copilot ([EPC Group](https://www.epcgroup.net/blog/microsoft-365-copilot-pricing-licensing-enterprise-guide-2026), [Teach AI Tools](https://teachaitools.blog/blog/microsoft-agent-365-and-e7-frontier-suite-launch-may-2026-what-the-99-tier-actually-changes)). But even the $99 seat does **not** include agent consumption — agent runs draw down credits (some sources describe "Agent Capacity Units") **billed separately**. Early-adopter data suggests a multi-step workflow costs $0.15–$0.25 and a complex orchestration $0.40–$0.80 per run; at 100,000 agent actions/month across 1,000 users, **variable costs run $1,500–$25,000/month on top of the seats** ([Teach AI Tools](https://teachaitools.blog/blog/microsoft-agent-365-and-e7-frontier-suite-launch-may-2026-what-the-99-tier-actually-changes)).

### 2.2 OpenAI: credits arrive inside ChatGPT

OpenAI has now brought consumption metering *inside* the ChatGPT Business/Enterprise product — not just its developer API. As of 2026, its **ChatGPT rate card** meters premium features in **credits**:

| Feature | Unit | Credits |
|---|---|---|
| **Agent mode** | 1 message | **30** |
| **Deep research** | 1 task | **50** |
| **Images** | 1 generation | **5** |
| **Voice** | 1 minute | **5** |

*Source: [OpenAI ChatGPT Rate Card](https://help.openai.com/en/articles/11481834-chatgpt-rate-card-business-enterpriseedu).*

More significantly, OpenAI has moved its **Workspace Agents, ChatGPT for Excel/Sheets, and ChatGPT for PowerPoint** to **token-based credit pricing** — e.g., GPT-5.5 at **125 credits per 1M input, 12.5 per 1M cached input, 750 per 1M output**, with GPT-5.4 at exactly half that ([OpenAI Rate Card](https://help.openai.com/en/articles/11481834-chatgpt-rate-card-business-enterpriseedu)). On **April 2, 2026**, OpenAI also converted Codex from per-message to **API-token-aligned pricing** and introduced **Codex-only seats with no fixed fee** — pure pay-as-you-go, billed on token consumption ([OpenAI Help](https://help.openai.com/en/articles/8792536-managing-billing-and-seats-in-chatgpt-business)).

The pattern: OpenAI *cut* the flat Business seat by $5 while simultaneously moving the high-cost features (agents, deep research, code) onto meters. **The flat part gets cheaper; the variable part grows.** Expect this to continue.

### 2.3 Anthropic: seat + usage, tokens unbundled

Anthropic's public Enterprise pricing is stated as **"Seat price + usage at API rates"** with admin-set org spend limits ([Claude pricing](https://claude.com/pricing)). Its **April 2026 removal of bundled tokens** from enterprise seat deals ([The Register](https://www.theregister.com/software/2026/04/16/anthropic-ejects-bundled-tokens-from-enterprise-seat-deal/5226555)) confirms the trajectory. Anthropic has also introduced **Managed Agents** billed at **$0.08 per session-hour of runtime plus standard token costs** ([MetaCTO](https://www.metacto.com/blogs/anthropic-api-pricing-a-full-breakdown-of-costs-and-integration)) — an explicit *time-based* meter for agents, the clearest acknowledgment yet that agent cost is a function of how long something runs, not just how much it says.

### 2.4 The macro signal

This is not three vendors making independent choices; it's an industry-wide repricing. Gartner projects worldwide AI spending of **~$2.59 trillion in 2026, up 47%**, with **AI agent software alone at $206.5B, rising to $376.3B in 2027** ([Business Age](https://bizage.jp/en/articles/gartner-ai-spending-2026-agent-economy)). Gartner forecasts **agentic AI spending will overtake chatbot/assistant spending by 2027** ([Software Strategies Blog](https://softwarestrategiesblog.com/2026/02/16/gartner-forecasts-agentic-ai-overtakes-chatbot-spending-2027/)). Translation for the boardroom: **the cheap, predictable, per-seat era of AI is the past; the metered, variable, agent era is the near future.** Budget accordingly.

---

## Part 3 — How to PLAN: budgeting, unit economics, and procurement

### 3.1 Budget in unit economics, not totals

The single most important discipline: **stop tracking total spend and start tracking unit cost.** A total ("we spent $80K on AI last month") tells you nothing about whether AI is working. A unit cost — **cost per employee, cost per task, cost per resolved ticket** — tells you whether the economics scale. FinOps practitioners frame it as expressing spend as *unit cost per 1k inferences / per 1k tokens / per task, tracked as a trend, not a total* ([ASleekGeek](https://asleekgeek.com/articles/finops-for-ai)).

Three unit metrics to instrument from day one:
- **Cost per active employee** — total AI spend ÷ people actually using it (not licensed). Reveals shelfware.
- **Cost per task/agent run** — for any automated workflow, the fully-loaded cost of one completed job.
- **Cost per business outcome** — cost per resolved support ticket, per qualified lead, per closed book. This is the number that survives a CFO review.

### 3.2 The pilot-to-scale cost curve (where budgets die)

The most dangerous moment is the transition from pilot to production, because **usage does not scale linearly with rollout — it jumps.** Successful pilots routinely drive a **5–10× jump in query volume within the first month of production** ([Hiren Dhaduk / LinkedIn](https://www.linkedin.com/posts/hiren-dhaduk_youre-locking-in-2026-cloud-waste-right-activity-7396888212432101376-tG4j)). A pilot that cost $2,000/month can become a $15,000/month production system not because anything broke, but because it worked.

A defensible staged budget (percent of first-year AI program spend):

| Phase | Share of budget | What it funds |
|---|---|---|
| Discovery | 5–15% | Readiness, use-case selection |
| Pilot | 10–25% | Funded 90-day quick-starts with KPI gates |
| Scale | 30–50% | Data engineering, integration, heavy consumption |
| Operate (recurring) | 15–30% | Monitoring, governance, retraining |

*Source: [Adoptify budget guide](https://www.adoptify.ai/blogs/2026-ai-adoption-cost-guide-budget-breakdown-for-enterprises/).* Practitioners also recommend **reserving 10–25% of budget as contingency for consumption spikes** ([Adoptify](https://www.adoptify.ai/blogs/2026-ai-adoption-cost-guide-budget-breakdown-for-enterprises/)). Set two hard tripwires before any production rollout: a **cost-per-query ceiling** (investigate when it exceeds pilot baseline by 20%) and a **daily spend alert at 3× average pilot-phase daily cost** ([Hiren Dhaduk](https://www.linkedin.com/posts/hiren-dhaduk_youre-locking-in-2026-cloud-waste-right-activity-7396888212432101376-tG4j)).

### 3.3 Showback → chargeback (making teams own their AI cost)

Because AI cost is now variable and driven by behavior, the organization needs a way to attribute it. The FinOps discipline offers a well-worn ladder:

- **Showback** — you *report* each team's AI consumption without moving it onto their budget. This builds cost awareness cheaply and is the right first step ([Mavvrik](https://www.mavvrik.ai/blog/chargeback-vs-showback/)).
- **Chargeback** — you *bill* the consuming business unit's P&L, creating genuine financial accountability ([FinOps Foundation](https://www.finops.org/framework/capabilities/invoicing-chargeback/)).

Start with showback in the pilot/scale phase; graduate the highest-consuming teams to chargeback once usage stabilizes. The behavioral effect is large: teams that see their own token bill trim their own prompts.

### 3.4 Procurement and licensing strategy

**Business vs. Enterprise — when to jump.** The self-service Business tiers (ChatGPT Business at ~$20, Claude Team at ~$20, Copilot Business at $21) are the right proving grounds — low minimums (2–5 seats), no long negotiation. Move to Enterprise only when you need the things Enterprise actually adds: **SSO/SCIM, audit logs, data-residency, admin spend limits, no-training guarantees, and negotiated volume discounts.** OpenAI Enterprise carries a **150-seat minimum and ~$108K/year floor** ([Beam Cloud](https://www.beam.cloud/blog/chatgpt-enterprise-pricing)); don't cross that line for features you won't use.

**Volume changes the math.** OpenAI Enterprise reportedly falls toward **~$40/user at 5,000+ seats** vs. ~$60 at the 150-seat floor ([Beam Cloud](https://www.beam.cloud/blog/chatgpt-enterprise-pricing)). Negotiate on seat count and term.

**For the credit/token meters, buy in the right order.** For any consumption-metered service, the universal pattern is: **start on pay-as-you-go during the pilot** (no commitment, you learn your real run-rate), then **switch to prepaid packs or committed volume once usage is predictable** (Copilot packs are ~20% cheaper per credit than PAYG — $0.008 vs. $0.01) ([Reveal Compliance](https://revealcompliance.com/blog/copilot-studio-pricing)). Committing to capacity before you know your run-rate is how organizations end up with unused prepaid packs that don't roll over.

**EA vs. CSP.** Large enterprises with an Enterprise Agreement get the deepest discounts and consolidated billing but least flexibility; the CSP (partner/reseller) channel offers month-to-month agility and is better for pilots and fluctuating agent workloads. Many enterprises run **both**: EA for the stable per-seat core, CSP/Azure PAYG for the variable agent meter.

**Watch the batch and caching discounts** — these are procurement levers, not just engineering ones. Anthropic's **Batch API is 50% off** for non-urgent work, and **prompt caching cuts cached input up to 90%** ([Claude Platform Docs](https://platform.claude.com/docs/en/about-claude/pricing)). For any workload that isn't real-time (overnight reports, bulk document processing), insisting on batch pricing halves the bill.

---

## Part 4 — Strategies across the four escalating tiers

Cost behavior changes fundamentally as you move up the ladder from ad-hoc chat to autonomous 24/7 systems. Each tier has a different dominant cost driver and a different set of control levers.

### Tier A — Chats (ad-hoc conversational use)

**What it is:** an employee opening ChatGPT or Copilot and asking questions.
**Cost model:** per-seat. Flat, capped, predictable.
**Dominant driver:** *seats* (and specifically, *unused* seats).
**The real risk here is shelfware, not overspend.** The MIT finding that 95% of pilots show no ROI ([Fortune](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/)) is largely a Tier-A adoption problem — companies buy thousands of seats, a fraction get used.
**Control levers:**
- Track **active vs. licensed** seats monthly; reclaim idle licenses.
- Use the cheaper Business tier for evaluation before Enterprise commitments.
- Since this tier is capped, spend your governance energy on *adoption*, not *cost control*.

### Tier B — Assistants (embedded copilots in the workflow)

**What it is:** Copilot inside Word/Excel/Teams; an assistant embedded in your CRM.
**Cost model:** still mostly per-seat, but the *hybrid meter starts showing up* here — premium features (deep research, image generation, workspace agents) begin drawing credits (see 2.2).
**Dominant driver:** *model choice* and *context length* — because assistants automatically pull in documents and history, inflating input tokens on every call.
**Control levers:**
- **Understand which features are zero-rated vs. metered.** A licensed Copilot user gets basic agent actions free but pays credits beyond that ([A Guide to Cloud](https://www.aguidetocloud.com/blog/copilot-credits-explained/)). Educate power users.
- **Right-size the default model.** Most assistant tasks (summaries, drafting, extraction) run fine on a mid or small model at 5–15% of frontier cost ([Degenito](https://degenito.ai/blog/best-practices-for-ai-agent-finops-control-llm-spend-at-scale/)).

### Tier C — Agents (task automation)

**What it is:** a bot that completes a multi-step task — triages a ticket, reconciles an invoice, drafts and sends an email.
**Cost model:** **fully consumption-based.** This is where you leave per-seat behind. Each run costs credits/tokens; volume drives the bill.
**Dominant driver:** *tokens × runs* — and crucially, the *number of steps in the agent's loop*.
**Why cost jumps here:** an agent doesn't make one model call; it makes a *loop* of calls — think, act, observe, think again. A single task can trigger a dozen model calls. Microsoft's rate card makes this concrete: a data-grounded agent turn runs 12–22 credits, and a complex orchestration can hit 40–80 credits (~$0.40–$0.80) *per run* ([Teach AI Tools](https://teachaitools.blog/blog/microsoft-agent-365-and-e7-frontier-suite-launch-may-2026-what-the-99-tier-actually-changes)).

**Control levers (the FinOps stack).** Teams combining these report **47–85% cost reductions without quality loss** ([Ruh AI](https://www.ruh.ai/blogs/agent-cost-optimization-playbook-ai-employees-smaller-bills)):

1. **Model routing / tiering** — send easy tasks to cheap models, escalate only the hard 10% to the flagship. The recommended distribution is **70% budget model / 20% mid-tier / 10% premium**, which cuts average per-query cost **60–80%** vs. routing everything through a flagship ([Ruh AI](https://www.ruh.ai/blogs/agent-cost-optimization-playbook-ai-employees-smaller-bills), [Premai via Ruh](https://www.ruh.ai/blogs/agent-cost-optimization-playbook-ai-employees-smaller-bills)).
2. **Prompt caching** — reuse the fixed parts of a prompt (system instructions, tool definitions) instead of re-sending them every call. Anthropic and OpenAI both now discount cached input **~90%**; caching reduces API costs **45–80%** ([Redis via Ruh AI](https://www.ruh.ai/blogs/agent-cost-optimization-playbook-ai-employees-smaller-bills), [Ofox](https://ofox.ai/blog/prompt-caching-cost-math-anthropic-vs-openai-2026/)).
3. **Semantic caching** — return a cached answer for a *similar* (not identical) question at near-zero cost.
4. **Budget guardrails** — hard per-agent, per-team, per-project spend caps: alert at 80%, throttle or suspend at 100% ([MeshAI](https://meshai.dev/blog/ai-agent-cost-optimization)).
5. **Circuit breakers** — a hard limit (~10) on how many steps an agent may take per task, so a stuck reasoning loop can't burn dozens of calls on one request ([Arun Baby](https://www.arunbaby.com/ai-agents/0047-cost-management-for-agents/)).
6. **Batch processing** — route non-urgent agent work through the 50%-off Batch API.

**A candid caveat leaders should hear:** these levers are not free wins you can bolt on blindly. One rigorous test found that naively adding caching, routing, and a budget cap actually made the bill *go up* — the SDK was already caching, and routing "taxed every ticket to save on three." The lesson: **cost control is measurement, not lever-collecting** ([Cost Control for AI Agents, EP08](https://www.youtube.com/watch?v=qNjuZtvnld4)). Instrument first; optimize what the data proves is expensive.

### Tier D — Agentic systems, including 24/7 autonomous operation

**What it is:** a fleet of agents running continuously — monitoring, responding, orchestrating other agents — without a human pressing "go."
**Cost model:** consumption-based *and time-based*. This is the tier where cost scales non-linearly and where the biggest surprises live.
**Dominant driver:** *runtime and held state* — not tokens per se, but **hours the system is alive.**

**This deserves the deepest analysis, because it inverts the cost logic executives are used to.** For Tiers A–C, cost is a function of *work requested*. For always-on systems, a large share of cost is a function of *time elapsed while doing nothing useful*. Three specific, non-obvious dynamics:

**1. Idle compute — the line item nobody puts on the invoice.** A long-running agent spends much of its life *waiting* — for a slow API, a queued tool, or a human who "went to lunch" while the meter keeps ticking. As one analysis puts it, the shift is *"from cost per call to cost per hour of held state"* — that is the entire story of long-running agents. A customer's job can sit blocked on a human for forty minutes while resources stay warm and billed ([GaaS](https://gaas.co.com/economics/the-economics-of-long-running-agents-hours-not/)). Anthropic's own move to bill Managed Agents at **$0.08 per session-hour** ([MetaCTO](https://www.metacto.com/blogs/anthropic-api-pricing-a-full-breakdown-of-costs-and-integration)) is the vendor formalizing exactly this: *you pay for the clock, not just the words.*

**2. Wakeup storms and runaway loops.** An autonomous agent that polls or retries can enter a pathological loop — repeatedly firing model calls with no progress. Without a hard cap, this compounds silently. Academic work on continuous-operation architecture (the "Heart" three-chamber design) reports running **75 agents continuously, 315 million heartbeats per year, at $0 token cost for keep-alive** — by separating cheap "liveness" signals from expensive "intelligence" calls, hard-capping daily model invocations with an atomic counter, and using a circuit breaker to stop wakeup storms ([Heart architecture](https://www.youtube.com/watch?v=qcisF1uFBe0)). The engineering lesson generalizes: **the expensive part of an agent is the loop and the orchestration, not the individual model call** ([Anthropic guidance, via GaaS](https://gaas.co.com/economics/the-economics-of-long-running-agents-hours-not/)).

**3. The inference-at-scale cost shift.** Once systems run 24/7 across thousands or millions of transactions, the dominant cost center moves from one-time *build* to ongoing *operation*. Industry analysis warns companies are **underestimating total AI cost by 30%+**, with continuous inference and monitoring now exceeding original development cost ([GlobeNewswire / Ramsey Theory](https://www.globenewswire.com/news-release/2026/04/02/3267325/0/en/dan-herbatschek-sees-1-trillion-ai-spend-crisis-as-enterprise-ai-costs-surge.html)). Gartner has even warned that **AI coding agents will cost more than the real developers** they were meant to augment ([Computer Weekly](https://www.computerweekly.com/news/366645054/Gartner-AI-coding-agents-will-cost-more-than-real-developers)).

**Control levers specific to Tier D (in addition to all of Tier C's):**

- **Separate liveness from intelligence.** Keep the "is it alive?" heartbeat on a cheap, non-LLM channel; only invoke the expensive model when there's genuine work. This alone eliminated keep-alive token cost entirely in the Heart deployment ([Heart architecture](https://www.youtube.com/watch?v=qcisF1uFBe0)).
- **Hard daily invocation caps** with atomic counters — a fixed ceiling on how many times any agent can wake the model per day.
- **"Nothing to do" fast-exit** — let an agent that wakes to an empty queue immediately roll back with zero cost rather than "thinking about" the emptiness.
- **Reserve-commit budgeting where possible.** A known gap in current tools: they track cost *after* the call completes, so a runaway response has already been paid for by the time you see it ([Cycles](https://runcycles.io/blog/ai-agent-cost-control-2026-litellm-helicone-openrouter-runtime-authority)). Where your platform supports pre-authorizing a budget before an agent runs, use it.
- **Kill failure loops early** — timeouts, retry caps, and canary gating when a workflow's success rate drops below ~85% ([HireNinja](https://blog.hireninja.com/2025/11/21/ai-agent-finops-a-30-day-playbook-to-cut-costs-25-40-with-opentelemetry-and-smart-model-routing/)).
- **Right-size to time-based pricing.** For agents that genuinely need to run continuously, evaluate time-metered offerings ($0.08/session-hour) against pure token pricing — for a mostly-idle 24/7 agent, per-hour can be cheaper than per-token, and vice versa. This is a real procurement decision, not just an engineering one.

### The four-tier summary

| Tier | Cost model | Dominant driver | #1 control lever | #1 risk |
|---|---|---|---|---|
| **A — Chats** | Per-seat (flat) | Seats | Reclaim idle seats | Shelfware, not overspend |
| **B — Assistants** | Seat + emerging credits | Model + context | Right-size default model | Silent feature metering |
| **C — Agents** | Consumption | Tokens × steps | Model routing (70/20/10) + circuit breakers | Non-linear volume jump |
| **D — Agentic 24/7** | Consumption + time | Runtime / held state | Separate liveness from intelligence + daily caps | Idle compute, wakeup storms |

---

## Part 5 — A 90-day action plan for the executive team

1. **Instrument before you scale.** Deploy usage/cost telemetry from day one. You cannot manage a variable meter you can't see ([HireNinja](https://blog.hireninja.com/2025/11/21/ai-agent-finops-a-30-day-playbook-to-cut-costs-25-40-with-opentelemetry-and-smart-model-routing/)).
2. **Define your three unit metrics** — cost per active user, per task, per business outcome — and report them monthly, not quarterly.
3. **Set tripwires now:** a cost-per-query ceiling (+20% over pilot) and a daily spend alert (3× pilot average).
4. **Assign a FinOps owner for AI** who reforecasts monthly. This is a named human, not a committee.
5. **Map your Copilot/Enterprise licensing maze** — compute true *base + add-on* cost, and account for the July 2026 base-price increases before renewal.
6. **Buy meters PAYG first, commit later.** Never prepay capacity before you know your run-rate.
7. **Before deploying any 24/7 agent, require three things in the design:** a hard step/circuit-breaker cap, a daily invocation ceiling, and a "nothing to do" fast-exit. Treat an always-on agent without a budget cap as a financial incident waiting to happen.
8. **Govern with budgets, not bans.** The goal isn't to slow AI adoption — it's to make its cost visible, attributable, and capped so that adoption can accelerate safely.

---

### Bottom line for the board

AI is migrating enterprises from a **seat economy** (linear, predictable, human-driven) to a **consumption economy** (non-linear, variable, machine-driven). All three major vendors — Microsoft, OpenAI, Anthropic — are converging on the same hybrid: a flat seat for people, a separate meter for machine work. The flat part is getting cheaper; the metered part is where your future budget risk lives. The organizations that win won't be the ones that spend the least on AI — they'll be the ones that can **measure cost per outcome, attribute it to the teams creating it, and cap the autonomous systems that would otherwise scale cost faster than value.**

*The single highest-leverage governance rule: **no autonomous agent goes to production without a hard budget cap and a circuit breaker.** Everything else is optimization; this is the seatbelt.*
