Weekly briefing

Budgets grew kill switches. Savings claims got metered.

Inside: Atlassian's per-engineer AI wallets, the token saver that billed 7.6% more while reporting 96.2M tokens saved, McKinsey's finding that a fifth or more of AI spend is unattributed, and the disappearance of every Anthropic and Google model from OpenRouter's top twelve.

August 3, 20266 source-linked reads
Editor's note

The last issue argued the fight was no longer about the size of the bill but about custody of the instrument. The answer arrived faster than expected, and it is smaller than a platform: the meter is being handed to the individual engineer with a ceiling attached. Atlassian's internal memo, obtained by Guardian Australia, sets monthly AI wallets between $500 and $2,000 across four tools and pauses usage at zero. PMG caps each user at $50 a day. U.S. Bank's chief AI officer describes killing timeboxed experiments before they spend rather than after. None of these is a leaderboard, and that is the point — a leaderboard rewards the number going up, and a wallet makes someone own it going down.

The complication is that the field started enforcing before it finished measuring. The sharpest item below is not a manifesto but a benchmark. A token-saving proxy advertising 60 to 90% savings went through 425 billed trials and came out 7.6% more expensive at low effort, having reported 96.2 million tokens saved on its own dashboard the entire time. That gap between a savings counter and an invoice is the thing to hold while reading everything else. McKinsey puts 20 to 30% of enterprise AI spend in the unattributed column, which means most caps are being wired to a number nobody has reconciled. A kill switch on an unaudited meter does not give you control. It gives you a faster way to act on a mistake.

Top stories

What mattered this week

Walletsthe Guardian

The cap arrived, and it has an engineer's name on it.

Guardian Australia saw the internal memo: Atlassian gives R&D staff monthly AI wallets ranging from $500 to $2,000 across four tools including Claude Code, fires alerts as the ceiling approaches, and stops usage at zero. So far every request for more has been granted, which makes this a budget with a tripwire rather than rationing. For scale, Elastic's ANZ manager Jeremy Pell puts the share of Australian organisations capping agent token or API use at 9%.

Takeaway: Gartner's Arun Chandrasekaran supplies the reason the ceiling has to sit on the person: per-token prices keep falling, yet a single agent can launch further agents and compose its own prompts, so a ceiling defined per model or per call gets outrun by the very thing it was meant to bound. One caution before borrowing the policy — Atlassian never ran a tokenmaxxing programme, so this is a company that skipped the phase describing its defaults, not a company reversing course.

Read source note
MeteringTech Times

A proxy advertising 60 to 90% savings billed 7.6% more.

A JetBrains benchmark ran 'rtk' — a proxy marketed on cutting Claude Code token use by 60 to 90% — across 425 billed trials against a pinned Claude Code build and Sonnet 5. At low effort the median session came out 7.6% more expensive. Uncached input, the only thing rtk actually compresses, moved 3.2% and never reached significance, while turns rose 13.8% and cache reads 14.3%. Output quality did not move to pay for the difference either — 71 of the 80 low-effort tasks finished tied. Throughout all of it, the tool's own dashboard was reporting 96.2 million tokens saved.

Takeaway: Those two numbers cannot both be true against the same bill, and the resolution is that a savings counter scores itself against a run that never happened. The reusable part is the method, not the verdict: researcher Denis Shiryaev predicted the roughly 3% ceiling for about zero dollars by replaying 83 saved transcripts before paying for a single trial. Replay before you buy, and grade every optimizer on an invoice rather than on its own dashboard.

Read source note
Field readAP News

The fad is buckling, and the arithmetic is why.

AP reports the tokenmaxxing fad giving way as costs climb without matching productivity, anchored on a new Moody's report whose AI analytics head Vincent Gusdorf says the bills piled up before teams learned disciplined use. The figure that travels furthest belongs to Bain's Jue Wang, who puts the doubling interval on client token bills at roughly every second month. Run $200 a head monthly across a 20,000-strong engineering org and you get a line item no general manager wrote down in advance.

Takeaway: Wang diagnoses a routing failure rather than an appetite problem — teams reaching for the most expensive model to draft routine email — which makes the fix a triage rule instead of a smaller ambition. Keep Mozilla CTO Raffi Krikorian's framing nearby: this is the new lines-of-code metric, which is a prediction that it gets gamed the moment anyone is rewarded for it.

Read source note
FinOpsMcKinsey & Company

Nine in ten enterprises are already past their AI budget.

McKinsey's report finds 93% of enterprises overrunning AI budgets, with spend rising nearly fourfold as pilots go company-wide and the same task costing up to 30 times the tokens depending on how it is run. Somewhere between a fifth and a third of that spend traces back to nothing at all, and only a fifth to a quarter of firms have mature AI FinOps. The report also records the vibe shift plainly: the word tokenmaxxing has turned pejorative inside enterprises, and some have quietly retired the AI user leaderboards they were promoting a year ago.

Takeaway: That unattributed share is the number governing every other control in this issue, because a cap on spend you cannot assign is a cap on the wrong person. And McKinsey's own lever list is duller and better than any ceiling: disciplined teams bank 20 to 30% across roughly 40 levers, with prompt caching alone cutting repeat input costs by up to about 90%. Fix the cache before you fix the culture.

Read source note
MutationGizmodo

Frontier-only mandates are a budget choice in a quality costume.

Reading the Wall Street Journal's CIO Journal, Gizmodo argues tokenmaxxing did not die so much as rebrand as frontier-only policy: Shopify reportedly barring engineers from cheaper models, and Olive founder Bill Nguyen burning 774 billion tokens in a single month at an estimated $4.5M. That works out near $5.80 per million tokens, which is a blended frontier rate — the signature of a shop that never routes anything down.

Takeaway: A rule forbidding cheap models reads as governance while quietly disabling the biggest lever on the bill that most teams still own. Send a third of those tokens down one tier and the identical workload sheds well over a million dollars inside a single billing period. Two cautions: the $4.5M is an estimate rather than an invoice, and the whole item is secondhand from the Journal's column, so pull the original before repeating either company's name as settled policy.

Read source note
Supply sideAnthropic

The half of the story nobody caps: cost per task fell.

Claude Opus 5 landed on July 24 carrying Opus 4.8's list price, and this is one vendor claim you can check against a third-party sheet — our own snapshot has both at $5 per million input and $25 output, against $10 and $50 for Fable 5, so 'near-Fable intelligence at half the price' holds on the price half at least. Every benchmark number is still Anthropic's own, though what it chooses to lead with is spend per completed task instead of a leaderboard score: on OSWorld 2.0 it claims to beat Fable 5's best mark for a little over a third of the money.

Takeaway: The effort setting is routing inside a single model, which makes it the cheapest routing available to anyone — no gateway, no migration, no cross-provider cache penalty. Anthropic reports passing more tasks than any rival even at its lowest effort. Treat that as a claim to test on your own workload, and run that test before committing a quarter to the multi-provider router that the cache arithmetic may never repay.

Read source note
Signals to watch

Where the next move is

Field readThe cap arrived and it is attached to a person, not a model. Atlassian issues R&D staff $500-$2,000 monthly AI wallets that pause at zero; PMG caps each user at $50 a day. A leaderboard rewards the number going up, a wallet makes someone own it going down.
Metering watchA proxy advertising 60-90% token savings billed 7.6% more across 425 trials while its own analytics reported 96.2M tokens saved. Any tool that scores itself against a counterfactual will always report a win — grade optimizers on the invoice.
Budget watchMcKinsey finds 93% of enterprises past their AI budget and 20-30% of spend unattributed, with only a fifth to a quarter running mature AI FinOps. Caps wired to unreconciled numbers just automate the error faster.
Routing watchFrontier-only mandates are the mutation to watch: Olive put 774 billion tokens through frontier models in a single month at a blended rate near $5.80 per million. Banning cheap models looks like governance and deletes the biggest lever most teams own.
SEO watchOur definition guide converts 0.92% from position 6.5 while the homepage converts 0.52% from 5.8 on the identical query. The better-ranking page is the wrong page, and it has stayed that way for two issues.
Infrastructure watch

Anthropic and Google are off the board entirely.

The rankings through August 2, pulled fresh for this issue, put DeepSeek V4 Flash first at about 817 billion tokens for the day, with a second DeepSeek V4 Flash build released July 31 close behind at 707 billion. Tencent's Hy3 is third, OpenAI's GPT 5.6 Luna fourth, Xiaomi's Mimo V2.5 fifth. Two weeks ago this board had Claude Opus 4.7 at ninth and Sonnet 5 at eleventh. Neither is in the top twelve now, and no Google model is either; of the three big American closed-model labs only OpenAI is still on it, though Nvidia and Poolside keep US-built models at sixth and eighth. Before reading any of that as share, note what the board actually meters: traffic that chose to pass through OpenRouter. The Claude Code and direct-API volume the rest of this issue is about never touches it. What it is a good census of is price-sensitive routing, and there the table has tilted hard — nine of the top twelve entries come from Chinese labs.

  • Sort by the trailing thirty days before calling anything a trend: Mimo V2.5 leads the month at 33.3 trillion tokens and Hy3 at 28.2 trillion, and neither of them led the day.
  • The July 31 DeepSeek build shows 707 billion tokens in a day against 1.3 trillion for the month. That shape is a launch ramp, not a position.
  • The price spread explains the standings: DeepSeek V4 Flash lists at $0.09 per million input against $5 for Claude Opus 5, so anything routed on price alone lands in the same handful of names.
Builder ecosystem

This fortnight the counters earned more than the spenders.

Sort the projects with real pull by what they do to a bill and they fall into two piles. LiteLLM, LangGraph, LlamaIndex and Qdrant spend tokens on your behalf; tiktoken, Langfuse, promptfoo and DSPy exist to tell you what happened afterwards. Given a fortnight in which a paid optimizer's own dashboard disagreed with the bill by more than the savings it promised, the counting half is where the leverage sat. The JetBrains method is copyable with nothing more than these: keep real sessions, replay them, count with a tokenizer you control, then compare against an invoice.

  • A gateway is only as honest as the counter behind it, so log tokens where you can recompute them independently.
  • Keep a corpus of replayable real sessions. It is the cheapest way to test an optimizer before it touches production spend.
  • Run cost and evals in the same pass, because savings that come back as retries show up in a different column than the one being celebrated.
Spend playbook

Audit the optimizer before you audit the engineers.

Most teams are running this sequence backwards: cap first, measure later. Invert it, and the first measurement is nearly free. Before paying for any token optimizer, replay a set of real transcripts through it offline and compute the ceiling on what it could possibly compress — in the JetBrains trial covered in this issue that ceiling was about 3% of input against an advertised 60 to 90%, and it was knowable for roughly nothing. Then grade the live result against an invoice rather than the tool's own counter, and watch the second-order costs a savings dashboard never shows: added turns, added cache reads, and retries after a quality regression.

  • Any tool reporting its own savings is scoring itself against a run that never happened. Grade it on the bill instead.
  • Compute your compressible fraction first. Nothing can save more than that, and it is almost always smaller than the pitch.
  • Track turns and cache reads next to tokens. Both rose here while 'tokens saved' climbed, and both cost real money.
Desk note

The page that ranks worse converts better, and we still have not acted.

Our own scoreboard, pulled live from Search Console through July 31, and it is not a flattering one. The definition queries we flagged last issue have not moved: 'what is tokenmaxxing' landing on the homepage draws 5,185 impressions at position 5.8 and converts 0.52%, against 0.51% at position 5.9 a fortnight ago. What is new is the cell sitting next to it. The same query landing on the guide written for it converts 0.92% from position 6.5 — worse rank, nearly double the click rate. The weakest cell on the board is the head term 'tokenmaxxing' at 4,855 impressions, 18 clicks, 0.37%, position 8.1. So the page that wins the ranking is the page that loses the click, and we have now watched that hold for two issues without shipping the fix.

  • Top SEO move this fortnight is the same one as last fortnight, which is itself the finding: route the definition queries to the guide instead of the homepage.
  • The case is no longer theoretical. The guide out-converts the homepage on identical intent from a worse position, so the better-ranking page is the wrong page.
  • Search Console runs a rolling ninety-day window, so compare rates and positions between issues rather than raw impression counts.

Read the token-spend tracking guide

Every control in this issue depends on a number you can recompute yourself. Here is how to build one — tokens by model, turns, retries, and cache behaviour — so the next savings claim has to get past a number of your own instead of your goodwill.

Continue reading
Issue links

Source notes from this issue

the Guardian source artwork
newsTG
news

Atlassian tightens tracking of staff AI use as other technology firms encourage ‘tokenmaxxing’

Guardian Australia saw an internal memo: Atlassian gave R&D staff monthly AI “wallets” of $500 to $2,000 spanning four tools including Claude Code. Alerts fire near the cap, usage pauses at zero, and no top-up has been refused yet.

tokenmaxxingexplainerworkplace-ai
Read note
Tech Times source artwork
newsTT
news

rtk Raises Claude Code Costs at Low Effort: JetBrains Benchmark Debunks 60–90% Claim

A JetBrains benchmark (July 20) ran ‘rtk,’ a proxy marketed to cut Claude Code tokens 60–90%, across 425 billed trials. At low effort it made sessions a median 7.6% MORE expensive—while rtk’s own analytics logged 96.2M tokens ‘saved.’

tokenmaxxingcoding-agentsagents
Read note
AP News source artwork
newsAN
news

Workplaces look for cheaper AI as ‘tokenmaxxing’ fades as a corporate fad

AP reports the tokenmaxxing fad is buckling as costs climb without matching productivity. Moody's AI analytics head Vincent Gusdorf, author of a new report, says bills piled in and teams realized the tools need disciplined use.

tokenmaxxingexplainerworkplace-ai
Read note
Generated Tokenmaxxing editorial thumbnail for The cost of intelligence: How CIOs can manage AI demand at scale - McKinsey & Company
newsM&
news

The cost of intelligence: How CIOs can manage AI demand at scale - McKinsey & Company

McKinsey’s July 20 report finds 93% of enterprises are already blowing past their AI budgets, with spend jumping nearly 4x as pilots go company-wide. The fix it prescribes: run “FinOps for AI” and treat tokens like cloud cost.

tokenmaxxingfinopsai-spend
Read note
Generated Tokenmaxxing editorial thumbnail for Tokenmaxxing Didn’t Die. It Mutated
newsG
newsmedium review

Tokenmaxxing Didn’t Die. It Mutated

Gizmodo reads the Wall Street Journal's CIO Journal and finds tokenmaxxing rebranded as frontier-only mandates: Shopify bars engineers from cheaper models, while Olive founder Bill Nguyen burned 774 billion tokens in a month.

tokenmaxxingexplainerworkplace-ai
Read note
Anthropic source artwork
newsA
news

Introducing Claude Opus 5

Anthropic shipped Claude Opus 5 on July 24 priced level with Opus 4.8, pitching near-Fable 5 intelligence at half the price. Its own Frontier-Bench v0.1 figures show more than double the predecessor's score, and cheaper per task.

tokenmaxxingcoding-agentsagents
Read note