Weekly briefing

The token count got fired from reviews and rehired by finance.

Meta pulls token counts out of performance reviews while Splunk, AWS and a new billing spec itemize every call. Plus: the DeepSeek ramp we flagged checked out.

September 21, 20266 source-linked reads
Editor's note

Goodhart's law got a live demonstration this month. Meta ranked engineers on a token leaderboard, found people picking up extra tasks just to climb it, and, according to a memo reported by The Information, has taken token counts out of performance reviews. Amazon ran into the same gaming and took its board down. If you wanted an obituary for the metric this site is named after, as a measure of people, that is roughly it.

The counting did not stop, though. It moved. Splunk has announced per-employee attribution of coding-agent spend, AWS will tell you which IAM principal made each Bedrock call, and FOCUS 1.5, due in December, is set to carry token-level usage and per-provider pricing in one billing format. Same number, new reader: as a score it gets gamed, as a cost signal it gets managed. Read each item below for who is looking at the count and what they are allowed to do with it.

Top stories

What mattered this week

IncentivesInfoWorld

Meta drops token counts from performance reviews

According to an internal memo from Maher Saba and Santosh Janardhan, adoption dashboards and token totals no longer feed ratings; managers are to judge how hard and how good the shipped work was. The trigger was the predictable one. Once the leaderboard counted, people took on work to feed it. This reaches us secondhand, InfoWorld relaying The Information, so treat the memo's wording as reported rather than read.

Takeaway: The part worth copying is what Meta kept. AI use is still tracked under a program from April, for training rather than ratings. That is the right split: the count survives as a signal about cost and adoption and stops being something anyone is paid to push up. A metric that sets pay gets optimized. One that only sets budgets mostly gets read.

Read source note
AttributionCisco Newsroom

Cisco's Splunk adds Tokenomics to track coding-agent token spend

Announced at Splunk .conf on September 15, the Tokenomics module inside Agent Observability splits token spend by agent, and by employee for coding-agent use, naming Claude Code, Codex and Cursor. That moves the question of who burned the budget into the same console that already watches how the agents behave.

Takeaway: Note the irony: a per-employee view of coding-agent spend is one sort order away from the leaderboard Meta has reportedly pulled out of reviews. What separates them is who gets the view, so settle before rollout whether it goes to budget owners or to line managers. The feature to test is the forecast, which Cisco says projects consumption before the billing period closes. It is a press release, with no pricing, accuracy figures or ship date for Tokenomics.

Read source note
StandardsIT Pro

From tokenmaxxing to valuemaxxing

IT Pro puts the what-replaces-the-leaderboard question to analysts at Gartner, IDC and 451 Research, plus HPE, and the answer on offer is outcome metrics. The Tokenomics Foundation's Mike Fuller objects that an outcome figure covers only half the ratio. The cost half is at least getting a standard: FOCUS 1.5, backed by the FinOps Foundation and due in December, carries token-level usage and per-provider pricing in one billing format.

Takeaway: Cost per outcome is cost divided by outcomes. Once FOCUS 1.5 lands, the numerator has a standard format and the denominator still has none, and most CIOs are told outcome tracking is more than a year out. Until then, Gartner's Stewart Buchanan offers the two cheapest habits in this issue: send deterministic jobs to plain rule-based code instead of a reasoning model, and take the first adequate answer instead of rerolling the prompt.

Read source note
Unit economicsExpress Computer

Enterprise AI budgets break at the handoff to production

New Relic India's Ganesh Narasimhadevara frames the paradox with one pair of numbers: blended cost per million tokens fell from $18.40 in Q1 2025 to $6.07 in Q1 2026, a two-thirds cut, and AI bills kept climbing anyway. The FinOps Foundation's 2026 survey, whose respondents manage roughly $83 billion of annual cloud spend, has 73% of organisations over their AI cost projections and only 43% with a formal AI governance policy.

Takeaway: Run the arithmetic the piece leaves implied. If the unit price is a third of what it was and the bill still grew, volume more than tripled. Cheaper tokens got spent, not banked. The firmest figures here are New Relic's own survey of 200 US technology leaders: 78% saw production incidents spike and 82% had a major failure tied to AI-written code within six months. That rework is the line nobody put in the forecast; the unattributed savings anecdotes are illustrations, not evidence.

Read source note
TelemetryAmazon Web Services

AWS maps three layers of Bedrock token-cost visibility

An AWS Cloud Financial Management post stacks Bedrock cost visibility in three layers: CloudWatch metrics alongside CUR 2.0 billing data, which now pins each call on the IAM principal that made it; a middle layer of model invocation logs; and client-side OpenTelemetry on top. The post dates from June, but it is still the most concrete build sheet in this issue for answering who made the call.

Takeaway: The 30 to 50% savings claim comes from routing, not from the telemetry itself. Moving routine work from Sonnet 4.5 ($3 in, $15 out per million tokens) to Haiku 4.5 ($1 and $5) is a flat two-thirds cut on both sides, before cached input at up to 90% off. Instrument the client, since that is where session cost and cache-hit ratios live, and note the uneven coverage: Claude Code gets the full worked example, while Cursor teams stop at layer two plus the vendor's own dashboard.

Read source note
OverheadXDA

Claude Code was using 51,000 tokens before I even typed a prompt — I fixed it

XDA's Mahnoor Faisal ran /context on a fresh session and found 51,400 tokens loaded before a single prompt. Built-in system tooling accounted for 28,500 of them, and nobody can remove that. Switching off four plugins installed for testing, plus auto-memory, brought the start down to roughly 41,400; custom agents alone fell from 943 tokens to 74.

Takeaway: Set aside the 28,500 fixed tokens and about 22,900 remained; one cleanup of plugins and auto-memory took roughly 44% of that. That preamble rides along on every turn: across a 40-turn session a 10,000-token trim is 400,000 fewer input tokens processed, whether they bill at the cached rate or not. Treat plugins like dependencies: review them, prune them, and change one at a time so you know which one paid.

Read source note
Signals to watch

Where the next move is

Reader demandDefinition intent still dominates. The spellings of “tokenmaxxing” account for most of our search impressions, and on the head term itself the guide written to answer it wins a click on under half a percent of its impressions.
Structural opportunityThe top item in this week's opportunity report is unchanged: rewrite the definition guide's title and description for click intent, and stop the homepage from answering the definition query on the guide's behalf.
Incentive watchMeta reportedly took token counts out of reviews and Amazon deleted its board, both after people gamed them. Any dashboard that ranks individuals by spend is a leaderboard, whatever the vendor calls it.
Agent watchStarting context is the overhead nobody budgets for. In XDA's test, about a fifth of a fresh session's starting context was plugins and memory features installed once and forgotten.
Infrastructure watchPer-call attribution is ordinary plumbing on Bedrock, Splunk has announced it for coding agents, and a billing standard is due in December. Outcome measurement is the half that still has no format.
Infrastructure watch

The DeepSeek ramp we flagged last week was real.

Last issue we said DeepSeek V4.1 Flash, then days old, was worth re-checking. On our OpenRouter snapshot pulled this morning, covering rankings through September 20, its trailing 30-day total has gone from 6.2 trillion tokens to 20.7 trillion. That is 14.5 trillion added across six days of data, or about 2.4 trillion a day, so its latest daily figure of 2.49 trillion is a run rate rather than a spike. It now sits second, a hair behind Z AI's GLM 5.3 Flash at 2.52 trillion. DeepSeek's July V4 Flash build slipped from 1.74 trillion a day to 1.12 trillion, though a Monday-to-Sunday comparison cannot separate migration from the weekend dip. Across its three builds on the board, DeepSeek's combined daily volume still rose, from 3.67 trillion to 4.03 trillion, on a Sunday.

  • Mind the weekday: this snapshot's source day is a Sunday and last week's was a Monday. GPT 5.6 Luna's daily count fell from 2.57 trillion to 1.21 trillion across that gap, yet its measured 30-day total rose to 49.5 trillion, effectively tied with DeepSeek V4 Flash for first.
  • Chinese labs hold 80% of top-twelve daily tokens on this Sunday snapshot. On the steadier 30-day totals the share is 75.4%, up from 73.4% a week ago.
  • Meta's Muse Spark 1.3 entered the top twelve at eleventh, and OpenAI's GPT 5.6 Sol dropped out.
  • This is one surface, not the market. Enterprises buying frontier models mostly call the vendor directly, and that traffic never passes through OpenRouter.
Builder ecosystem

The open-source attribution layer keeps growing, quietly.

Bedrock's CUR 2.0, and Splunk once Tokenomics ships, are the commercial route to knowing who spent what. The seventeen open-source projects the desk tracks are the do-it-yourself route, and this week's star counts say attention is concentrating rather than spreading: a gateway, an orchestrator and a tracer took the biggest gains while most of their peers barely moved.

  • LiteLLM added about 500 stars since last issue, the largest gain on our board, while fellow gateways Portkey, Helicone and TensorZero added 47, 9 and 2.
  • LangGraph and Langfuse added roughly 360 and 240. Orchestration plus tracing is the pairing that turns an agent run into something you can itemize.
  • OpenLLMetry, one open-source way to build the client-side OpenTelemetry layer AWS recommends, gained just 10. The plumbing that matters most rarely trends.
Spend playbook

Keep the meter. Drop the scoreboard.

Read together, the stories above separate two uses of the same count. The portable version is to attribute token spend precisely, show it to the people who own budgets, and keep it out of anything that decides pay or rank. Attribution is now plumbing you can buy or build; incentive design is still a choice you have to make on purpose.

  • Tag every call with a principal, whether a person, a service or an agent, before arguing about outcomes. On Bedrock, CUR 2.0 and client-side OpenTelemetry make this a configuration task.
  • Send per-person spend views to budget owners, not to performance ratings. When the question is cost, aggregate by team or workflow so nobody has a reason to perform for the chart.
  • Run /context, or your tool's equivalent, on a fresh session this week and write down the starting number. It is the cheapest audit in this issue.
  • Price the first adequate answer: cap rerolls, and move deterministic steps to ordinary code.
Desk note

Back on a seven-day review window, and one number that has not moved.

This is the first issue since the August outage built from an ordinary week of desk work. All six top stories were reviewed and published on the desk between September 16 and 19, and none has run in a previous briefing. The source pieces themselves are older than that window: four of the six predate this week, and the AWS post dates from June 29. Search Console, Ahrefs, PostHog and the router rankings all refreshed live this morning.

  • SEO watch, like for like: our definition guide's own row for “tokenmaxxing” shows 0.44% click-through at an average position of 7.1 on 12,859 impressions over the last 90 days. In June it was 0.42%. That is flat.
  • The flattering number sits right beside it. Across every page on the site, the query clicks through at 4.6% from position 2.5, because the homepage takes the click. The handover to the guide is what has not happened, and its title and description are the next fix.
  • The leaderboard queries are our standout click-through: “tokenmaxxing leaderboard” runs at 28.7% on 261 impressions, and its alternate spellings run near 30% on smaller volume. People still go looking for a tokenmaxxing leaderboard even as Meta reportedly takes token counts out of its reviews. Ours ranks models, not employees.
  • 210 candidates remain in the review queue.

Read tokenmaxxing vs AI outcomes

Build the meter without the scoreboard: what to count, who should see it, and how to put the accepted artifact next to the cost.

Continue reading
Issue links

Source notes from this issue

Generated Tokenmaxxing editorial thumbnail for Meta drops token counts from performance reviews
newsI
newsmedium review

Meta drops token counts from performance reviews

Meta's Maher Saba and Santosh Janardhan told staff in an internal memo that adoption dashboards and token counts are out of performance reviews; managers should weigh the difficulty and quality of shipped work instead.

tokenmaxxingworkplace-aimetrics
Read note
Cisco Newsroom source artwork
newsCN
news

Cisco's Splunk adds Tokenomics to track coding-agent token spend

At Splunk .conf on Sept. 15, Cisco added a Tokenomics module to Splunk Agent Observability. It attributes token spend across AI agents and across employees' use of coding agents, naming Claude Code, Codex and Cursor.

ai-spendcoding-agentsllm-observability
Read note
IT Pro source artwork
long-formIP
long-form

From tokenmaxxing to valuemaxxing

IT Pro canvasses Gartner, IDC, 451 Research and HPE on what replaces token leaderboards. The Tokenomics Foundation's Mike Fuller says the outcome-first 'valuemaxxing' fix only measures half the equation.

tokenmaxxingmetricscost-governance
Read note
Generated Tokenmaxxing editorial thumbnail for Enterprise AI budgets break at the handoff to production
long-formEC
long-form

Enterprise AI budgets break at the handoff to production

Express Computer interviews New Relic India's Ganesh Narasimhadevara on why AI bills keep climbing while the blended cost per million tokens has fallen over a year, from $18.40 in Q1 2025 down to $6.07 in Q1 2026.

tokenmaxxingai-spendcost-governance
Read note
Generated Tokenmaxxing editorial thumbnail for AWS maps three layers of Bedrock token-cost visibility
guideAW
guide

AWS maps three layers of Bedrock token-cost visibility

Rachanee Singprasong and Chitresh Saxena lay out three layers of Bedrock cost visibility: native CloudWatch metrics plus CUR 2.0 caller-identity attribution, then model invocation logging, then OpenTelemetry in the client.

ai-spendllm-observabilitycost-governance
Read note
XDA source artwork
long-formX
long-form

Claude Code was using 51,000 tokens before I even typed a prompt — I fixed it

Mahnoor Faisal opened a new Claude Code session, ran /context, and found 51,400 tokens already loaded. Disabling four test plugins and auto-memory got the starting context down to roughly 41,400 before any real prompt.

coding-agentstoken-consumptiontoken-waste
Read note