news

Workplaces look for cheaper AI as ‘tokenmaxxing’ fades as a corporate fad

AP reports the tokenmaxxing fad is buckling as costs climb without matching productivity. Moody's AI analytics head Vincent Gusdorf, author of a new report, says bills piled in and teams realized the tools need disciplined use.

Published 2026-07-28Source: AP News
AP News source artwork

Why it matters

Bain's Jue Wang says client token costs have been doubling almost every other month. At $200 per developer per month across 20,000 developers, that becomes a line item no general manager ever planned for.

Tokenmaxxing read

Wang's diagnosis is a routing failure: teams reach for Opus even to draft routine email. Her fix is model routing, sending easy queries to cheaper systems and hard tasks to powerful ones. Mozilla CTO Raffi Krikorian calls it the new lines-of-code metric.

Source takeaway

The backlash has teeth: Satya Nadella warns customers pay twice, in tokens and in proprietary data, and Palantir's Alex Karp says enterprises are privately livid about tokens that return nothing. Cheap Chinese open models cut the other way.

Topic links

Related projects

Tools that match this angle

#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

32.2K3.4KSource-available
tracesevalscosts
#5Direct
Evaluation

promptfoo

promptfoo/promptfoo

A CLI and CI workflow for testing prompts, agents, and RAG systems across models, with evals and red-team style checks.

23.8K2.1KMIT
prompt-evalscirag
#6In spirit
Evaluation

DSPy

stanfordnlp/dspy

A framework for programming and optimizing language-model pipelines rather than hand-tuning one prompt at a time.

36.5K3.1KMIT
optimizationprogrammingevals
Related feed

More source-linked context

Fortune source artwork
newsF
news

'You just hired a million bad employees': How the brief tokenmaxxing era delivered the opposite of what it promised

Fortune's Nick Lichtenberg argues the tokenmaxxing boom backfired: firms rolled out AI agents faster than they could manage them. Hebbia founder George Sivulka warns most companies effectively 'hired a million bad employees.'

tokenmaxxingexplainerworkplace-ai
Read note
theregister source artwork
newsT
news

Nvidia shows off Vera Rubin platform for tokenmaxxing

At an Nvidia lab briefing, Ian Buck pitched Vera Rubin as an AI-factory play: 10x the tokens per watt of a GB200 NVL72 in early CoreWeave runs on DeepSeek-R1, plus a Vera CPU said to double agent speed.

tokenmaxxingexplainerworkplace-ai
Read note
Generated Tokenmaxxing editorial thumbnail for Tokenmaxxing Didn’t Die. It Mutated
newsG
newsmedium review

Tokenmaxxing Didn’t Die. It Mutated

Gizmodo reads the Wall Street Journal's CIO Journal and finds tokenmaxxing rebranded as frontier-only mandates: Shopify bars engineers from cheaper models, while Olive founder Bill Nguyen burned 774 billion tokens in a month.

tokenmaxxingexplainerworkplace-ai
Read note