Weekly briefing

Everyone is building a meter. The fight is over who owns it.

Inside: why Palantir wants per-token pricing dead, the earnings miss Chamath sees coming, a disputed $1.7M in a $34M token audit, and the cache math that makes cross-provider routing backfire.

July 20, 20266 source-linked reads
Editor's note

Two weeks ago the field had just proved that spend discipline works. This stretch it turned on the instrument itself. Palantir's manifesto does not ask you to use fewer tokens — it asks why anyone is billing by the token at all. O'Reilly's Mike Loukides reaches the same place from the opposite direction: once access is metered and each generation costs more per token, performative token-burning is simply the easiest thing to strike from a budget.

The through-line worth holding is that a meter is never neutral. Whoever owns it sets the defaults, wins the arguments, and books the margin. That is why this week's product news matters more than the manifestos: Snowflake shipped per-user token quotas into the data platform and Citrix pushed token accounting down into the network path that every request traverses anyway. Both move the number off the vendor's console. And both run into the same complication — if cross-provider routing wrecks your prompt cache, the cheapest path is to stay inside one vendor, and the meter goes right back where it started. Read every item below as a claim about custody of the number, not just the size of the bill.

Top stories

What mattered this week

VerdictO’Reilly Media

O'Reilly calls the tokenmaxxing era over — on pricing, not principle.

Mike Loukides argues the trend dies the moment finance reads an invoice, and his evidence is a price list: GitHub Copilot retired unlimited access in favor of penny-denominated credits, GPT-5.5 arrived at double the per-token cost of GPT-5.4, and Claude Fable at double Claude Opus 4.8. Nobody had to win the argument about waste. The meter settled it.

Takeaway: Treat this as a schedule, not an opinion. Each model generation is landing at roughly twice the last one's per-token price, so the optimization you postpone now gets priced in the moment you upgrade — and the move is to delete waste, not to ration usage.

Read source note
ManifestoThe Next Web

Palantir's answer to tokenmaxxing is to stop paying by the token.

A nine-point 'AI sovereignty' post brands tokenmaxxing a hit of false progress and points the gun at metered pricing itself, with Alex Karp asking why the labs charge for tokens at all. It escalates the debate from a spending habit to a business model worth escaping.

Takeaway: Discount the messenger, keep the question. Telling enterprises to hold their own weights doubles as marketing for an air-gapped stack, and The Next Web notes the sovereignty logic boomeranged when France's foreign intelligence service dropped Palantir for a homegrown rival. The durable fact underneath: aggregate AI spend keeps rising while the unit price falls, which means volume is eating the discount.

Read source note
EarningsCNBC

Chamath thinks the bill shows up first as a bad quarter.

Palihapitiya told CNBC that token spend is largely invisible to the chief executives and finance chiefs who answer for it, and that the first visible symptom will be an EPS miss of a few pennies nobody can account for. His own 8090 is on pace to spend more than $10M a year on AI alone — a figure he calls scary. The problem he names is not the size of the spend but that it never rolls up to anyone.

Takeaway: Unattributed spend is ungovernable spend. If a token bill cannot be mapped to a team and a P&L line, no cap will hold, because there is no one to hand the cap to — build the attribution before you build the budget.

Read source note
Meter noiseForbes

An audit of $34M in token spend flagged $1.7M as billing errors.

Forbes canvasses operators and finds them aligned that spend, not usage, is the number that matters now: watching deployment costs has gone from a rounding-error habit in 2023, at 3% of organizations, to standard practice at 58%. The sharpest detail sits mid-piece — Vaudit says an audit covering $34M of token spend flagged $1.7M in billing mistakes.

Takeaway: Five cents on the dollar in dispute, with two caveats: Vaudit sells token audits, so treat the rate as a marketing number, and nothing says the errors ran in your favor. The point is that the meter itself is noisy in both directions. Recount one month against your own logs and find out which way your gap runs before you spend a quarter tuning what the invoice claims.

Read source note
FinOpsSnowflake

Cost control stopped being a project and became a platform default.

Snowflake argues AI spend deserves the same treatment cloud bills eventually got, and backs the argument with product: quotas attached to individual users, budgets, and cost views that roll up to the org. The figure that travels further than the feature list comes from the FinOps Foundation, which now finds 98% of FinOps teams managing AI spend, against 31% two years back.

Takeaway: Read it as vendor framing and still take the signal. When quotas arrive as defaults, the cap lands on the engineer mid-task rather than on the invoice at month-end — decide deliberately which of your workloads should hit a wall versus a warning, because the platform will decide for you otherwise.

Read source note
Control planeERP Today

Citrix is moving token accounting into the network path.

ERP Today reports that NetScaler's new MCP Gateway steers agent traffic toward sanctioned servers and counts tokens on the way through, split by team, user, or application, and does it across competing model providers. GM Steve Shah's framing is that an agent's MCP call has become what the API call used to be — and anything sitting in that path sees enough to assign costs without needing a vendor's cooperation.

Takeaway: Watch the layer, not the launch. A provider-agnostic meter running inside infrastructure you already operate is what turns switching costs from a threat into a lever — and the tell for where this is heading is a private preview that parks a corporate chokepoint in front of Claude Code access.

Read source note
Signals to watch

Where the next move is

Field readThe debate moved from how much to spend to who owns the measurement. Palantir attacked per-token billing as a business model, and O'Reilly declared the era finished on pricing grounds alone.
Governance watchChamath Palihapitiya expects runaway token spend to surface as an unexplained EPS miss, because the bills never roll up to a team or a P&L. Attribution is the prerequisite for every cap.
Billing watchVaudit says it flagged $1.7M of billing errors in a $34M token audit — about five cents on the dollar, self-reported by a firm that sells audits. Either way, the meter is noisy in both directions.
Routing watchTechzine's counterargument: cross-provider routing triggers cold starts that discard prompt-cache savings, so provider-side routing wins and lock-in deepens. Price cache warmth before you price tokens.
SEO watch'What is tokenmaxxing' pulls 4,132 impressions at position 5.9 and converts 0.51% to clicks. The brand term improved to 1.63%; the definition queries never got the same treatment.
Infrastructure watch

The public board may be measuring the traffic that is about to move in-house.

OpenRouter's rankings through July 19 put Tencent's Hy3 first for the day near 1.56 trillion tokens, but Xiaomi's Mimo V2.5 still leads the trailing thirty days at about 25.1 trillion, with Deepseek V4 Flash close behind at 22.0 trillion — the day's leader and the month's leader are not the same model. Anthropic's Claude 4.7 Opus has slipped to ninth, from sixth two weeks ago, with Claude Sonnet 5 just behind at eleventh. Set that against Techzine's argument that every hop to a cheaper provider triggers a cold start and discards the prompt-cache savings routing was supposed to capture — the same piece notes Uber reportedly ran through a year of AI budget in roughly four months of Claude Code usage before anyone did this arithmetic. If that argument holds, the incentive is to consolidate onto one provider so the cache survives, provider-side routing wins, and cross-provider boards like this one stop describing where the work actually runs.

  • Sort by trailing-30-day tokens before calling anything a trend: Hy3 leads the day, Mimo leads the month, and only one of those is a pattern.
  • Kimi K3 entered at eighth for the day on roughly 154 billion tokens against a 380 billion thirty-day total — that gap is a days-old launch spiking, not a model winning.
  • Cache warmth is the routing variable nobody prices. Measure recomputation cost per hop before you route, or a cheaper model per token becomes a more expensive task.
Builder ecosystem

The tooling is converging on custody of the number.

The projects getting real adoption — gateways and routers like LiteLLM, orchestration like LangGraph, tracing like Langfuse, evals like promptfoo and DSPy, tokenizers and retrieval like tiktoken, LlamaIndex, and Qdrant — line up closely with the week's news. The gateway and tracing layers are the direct attempt to hold a copy of the meter outside the model vendor, the same move Citrix is making in the network and Snowflake in the warehouse. The rest are what keep the number honest once you have it.

  • A gateway turns model choice from a per-engineer habit into a policy you can audit after the fact.
  • Tracing is what makes a cost-reduction claim checkable by someone who does not trust you.
  • Your own tokenizer counts are the only way to catch the billing errors the Forbes piece documents.
Spend playbook

Reconcile the invoice before you optimize the usage.

FormAssembly CTO Bryan O'Neill, writing in InfoWorld, puts tokenmaxxing in a lineage with paying per line of code and story points — metrics that counted effort until teams learned to inflate them. His fix is to constrain intent rather than budget: decide what the work is before you authorize spend against it. His warning also applies to the meter this whole issue is arguing for, because a token counter is easier to game than a story point. That is why the number has to attach to an artifact rather than to a person.

  • Recount before you route. Reconciling a month of invoices against your own logs takes an afternoon; a routing project takes a quarter and may not survive the cache math.
  • Attach every agent run to an owner and an artifact, so the spend has somewhere to roll up before finance asks.
  • Bound retries and context growth before the run starts. A budget enforced after the fact is a report, not a control.
Desk note

We rank for the question and lose the click.

A transparency note on our own numbers: 'what is tokenmaxxing' now draws 4,132 impressions against 21 clicks — a 0.51% click rate at an average position of 5.9. Ranking sixth for a question and converting half a percent means the answer is visible in the result and there is no reason to click through. The headline term improved since the last issue: 'token maxxing' moved from 6,256 impressions at 1.2% to 9,303 at 1.63%, position 7.3 to 6.2. So the title work is landing on the brand term and has not touched the definition queries.

  • Top SEO move this week: the definition queries, not the brand term. 'What is tokenmaxxing' sits at position 5.9 and converts 0.51% of impressions — the weakest of the queries we track.
  • Those queries are still being answered by the homepage rather than the guide written for them, which is the kind of mismatch that ranks well and converts badly.
  • Every number above comes from a live pull through July 19, not a cached fallback.

Read the token-spend tracking guide

Every story this week is an argument about who holds the meter. Here is how to build your own copy of it — tokens by model, retries, and cache behavior — so the vendor's invoice is something you check rather than something you accept.

Continue reading
Issue links

Source notes from this issue

O’Reilly Radar: The End of Tokenmaxxing artwork
newsOM
news

The End of Tokenmaxxing

O'Reilly's Mike Loukides argues the tokenmaxxing era ends once finance notices the bill: GitHub Copilot swapped unlimited access for $0.01 credits, GPT-5.5 costs 2x GPT-5.4, and Claude Fable doubles Opus 4.8 per token.

tokenmaxxingexplainerworkplace-ai
Read note
Palantir AI sovereignty manifesto artwork
newsTN
newsmedium review

Palantir's 9-point manifesto decries tokenmaxxing and champions 'AI sovereignty'

Palantir dropped a 9-point 'AI sovereignty' manifesto on X, branding tokenmaxxing a hit of 'false progress' and taking direct aim at OpenAI and Anthropic's per-token pricing. CEO Alex Karp's jab: 'Why are they charging for tokens?'

tokenmaxxingexplainerworkplace-ai
Read note
Generated Tokenmaxxing editorial thumbnail for Chamath Palihapitiya says soaring AI token spend will hit company earnings
newsC
news

Chamath Palihapitiya says soaring AI token spend will hit company earnings

Investor Chamath Palihapitiya told CNBC that runaway AI token spend is largely invisible to CEOs and CFOs, and predicted it will eventually surface as an unexplained earnings miss of a few cents a share.

tokenmaxxingai-spendworkplace-ai
Read note
Generated Tokenmaxxing editorial thumbnail for After ‘Tokenmaxxing’, Token Spend Has Become The New Metric To Watch
newsF
news

After ‘Tokenmaxxing’, Token Spend Has Become The New Metric To Watch

Forbes contributor Tim Keary argues the tokenmaxxing push has flipped into cost discipline: with CFOs and boards watching, firms now track token spend per engineer and blend premium and cheaper models instead of maximizing usage.

tokenmaxxingcoding-agentsagents
Read note
Generated Tokenmaxxing editorial thumbnail for FinOps for AI: Snowflake's AI Cost Management and Governance Tools
newsS
news

FinOps for AI: Snowflake's AI Cost Management and Governance Tools

Snowflake's product team makes the case for 'FinOps for AI' — governing model spend the way cloud bills got governed — and rolls out per-user token quotas, budgets, and org-level cost views to meter Cortex and agent usage.

tokenmaxxingfinopsai-spend
Read note
Generated Tokenmaxxing editorial thumbnail for AI Agents Need a Gateway, and Citrix Is Putting NetScaler in the Middle
newsET
news

AI Agents Need a Gateway, and Citrix Is Putting NetScaler in the Middle

Citrix said on July 9 that NetScaler AI Gateway now carries an MCP Gateway, steering agent traffic to approved MCP servers while metering input and output tokens per team, user, or app across rival model providers.

tokenmaxxingagentstoken-consumption
Read note