news

‘Modelmaxxing’ Replaces ‘Tokenmaxxing’ for Firms Grappling With AI Costs

The Daily Upside charts enterprises trading token-hoarding for per-task model shopping. An IDC poll of 260 US decision-makers at firms above 1,000 staff found 47% already run a Chinese model somewhere, and 20% lean on them heavily.

Published 2026-07-26Source: The Daily Upside
The Daily Upside source artwork

Why it matters

Model choice is hardening into a procurement decision rather than an engineering preference. Counterpoint analyst Soumen Mandal expects standard workloads to commoditize, leaving how a model is deployed — not whose badge it carries — as the differentiator.

Tokenmaxxing read

Mandal puts the spread at under $5 per million output tokens for DeepSeek or Qwen against roughly $25–$30 for GPT 5.5 and Claude Opus. OpenRouter’s Adam Swick names the shape that harvests it: plan on the expensive model, hand the routine legs to cheap ones.

Source takeaway

Analyst-sourced and worth citing, but keep the two surveys straight: the 47% China-adoption number is IDC’s 260-person panel, while the 73% routing-adoption figure comes from a separate poll of roughly 100 respondents taken late last year.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

56.3K10.5KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

33K3.6KSource-available
tracesevalscosts
#10Direct
Routing

Portkey Gateway

Portkey-AI/gateway

An AI gateway for routing across LLMs with guardrails, provider abstraction, and an OpenAI-compatible API surface.

12.7K1.2KMIT
gatewayguardrailsrouting
Related feed

More source-linked context

Generated Tokenmaxxing editorial thumbnail for Amazon deletes devs’ tokenmaxxing leaderboard to minimize costs - InfoWorld
newsI
news

Amazon deletes devs’ tokenmaxxing leaderboard to minimize costs - InfoWorld

Amazon reportedly pulled an unofficial internal leaderboard that ranked employees by AI usage after it drove wasteful behavior and higher compute bills—workers started spinning up agents just to climb the rankings.

tokenmaxxingcost-governanceai-spend
Read note
Generated Tokenmaxxing editorial thumbnail for “Tokenmaxxing is real, expensive & it’s spreading”: AI budgets are exploding - The New Stack
newsTN
newsmedium review

“Tokenmaxxing is real, expensive & it’s spreading”: AI budgets are exploding - The New Stack

AI accountability startup Lanai debuted Token Tuner, a beta that scores each employee's efficiency by matching token usage and model choice to task complexity — peers burned 10x the tokens for half the efficiency in one beta.

ai-spendcost-governanceexplainer
Read note
Forbes source artwork
newsF
news

Companies With Goals Of AI Tokenmaxxing Are Foolishly Inspiring Employees To Waste Costly AI Resources

Forbes argues tokenmaxxing becomes a perverse incentive when companies set usage targets: employees learn to burn tokens, not to ship outcomes.

tokenmaxxingcost-governanceai-spend
Read note