news

The cost of intelligence: How CIOs can manage AI demand at scale - McKinsey & Company

McKinsey’s July 20 report finds 93% of enterprises are already blowing past their AI budgets, with spend jumping nearly 4x as pilots go company-wide. The fix it prescribes: run “FinOps for AI” and treat tokens like cloud cost.

Published 2026-07-20Source: McKinsey & Company
Generated Tokenmaxxing editorial thumbnail for The cost of intelligence: How CIOs can manage AI demand at scale - McKinsey & Company

Why it matters

The same task can burn up to 30x the tokens, and 20–30% of AI spend is simply unaccounted for. Only 20–25% of firms have mature AI FinOps, so most are flying blind—while a majority expect spend to climb another 25%+ over the next year.

Tokenmaxxing read

McKinsey names the vibe shift: ‘tokenmaxxing’ is now a dirty word, and some firms are pulling down their AI user leaderboards. Disciplined teams bank 20–30% across ~40 levers—prompt caching alone cuts repeat input costs up to ~90%. Govern the outcome, not the token.

Source takeaway

It reframes buy-vs-build as a live ‘buy, build, host, route, switch’ mix and pushes a central AI control plane as the single source of truth. Chargeback ties spend to the unit that caused it; forecasting maturity is worth ~10% more savings.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

55.2K10.2KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

32.2K3.5KSource-available
tracesevalscosts
#10Direct
Routing

Portkey Gateway

Portkey-AI/gateway

An AI gateway for routing across LLMs with guardrails, provider abstraction, and an OpenAI-compatible API surface.

12.6K1.2KMIT
gatewayguardrailsrouting
Related feed

More source-linked context

Generated Tokenmaxxing editorial thumbnail for FinOps for AI: Snowflake's AI Cost Management and Governance Tools
newsS
news

FinOps for AI: Snowflake's AI Cost Management and Governance Tools

Snowflake's product team makes the case for 'FinOps for AI' — governing model spend the way cloud bills got governed — and rolls out per-user token quotas, budgets, and org-level cost views to meter Cortex and agent usage.

tokenmaxxingfinopsai-spend
Read note
abhs.in — Abhishek Gautam source artwork
newsA—
newsmedium review

Kubernetes Becomes the AI Substrate: 66% of GenAI Inference, DRA GA, llm-d

A practitioner reading of June's CNCF news: 66% of orgs running GenAI inference do it on Kubernetes, DRA went GA, gang scheduling landed natively, and Nvidia and Google donated their DRA drivers — self-hosted inference is complete.

ai-spendcost-controlcost-governance
Read note
Procurement Magazine source artwork
newsPM
news

How Ramp is Fuelling AI Spend Management Expansion

Ramp closed a $750M round at a $44B valuation and is launching AI token spend management, procurement agents, and accounting agents on top of $1B+ annualized revenue and 70,000+ customers.

agentsai-spendcost-governance
Read note