news

The problem with AI model routing

Techzine’s Erik van Klinken argues cross-provider model routing can quietly backfire: each hop to a cheaper model triggers a cold start that throws away prompt-cache and context savings, so recomputation can cost more than routing saves.

Published 2026-07-06Source: Techzine Global
Generated Tokenmaxxing editorial thumbnail for The problem with AI model routing

Why it matters

Routing is pitched as the cure for runaway token bills, yet if it wrecks caching it can inflate them. With Anthropic’s inference margins reportedly 70% and $20/100/200 plans heavily subsidized, buyers chasing per-token savings may be tuning the wrong layer.

Tokenmaxxing read

The counterintuitive lever is caching, not routing. Van Klinken expects provider-side routing — staying inside one vendor to keep the cache warm — to win, deepening lock-in. Uber reportedly spent a full year’s AI budget within four months on Claude Code tokens before it clicked.

Source takeaway

Erik van Klinken (Techzine Global): a month ago buyers still reached for the largest model 95% of the time; with Fable 5 priced at twice per token of Opus 4.8, he bets vendors’ own routers, not third-party ones, capture the efficiency trade.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

56.5K10.7KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

33.2K3.6KSource-available
tracesevalscosts
#10Direct
Routing

Portkey Gateway

Portkey-AI/gateway

An AI gateway for routing across LLMs with guardrails, provider abstraction, and an OpenAI-compatible API surface.

12.7K1.2KMIT
gatewayguardrailsrouting
Related feed

More source-linked context

PYMNTS.com source artwork
newsP
news

AI Agents Just Got Their Own Company Credit Cards

Mercury launched Agent Cards through Mercury Spend: virtual cards an AI agent spends from inside company-set rules, with transactions outside them declined automatically and no way for the agent to raise its own limit.

tokenmaxxingagentsai-spend
Read note
HackerNoon source artwork
newsH
newsmedium review

Auto-Mode Routing: What Stop Us From Sending "How to Center a Div" to Claude 3.5 Opus | HackerNoon

The vCodeX team audited about 100,000 internal prompt logs, found roughly 65% were definitional questions or minor refactors hitting frontier endpoints by default, and built a sub-40ms complexity router across three model tiers.

tokenmaxxingcost-governanceai-spend
Read note
American Banker source artwork
newsAB
newsmedium review

Beyond token-maxing: How US Bank AI chief navigates costs

U.S. Bank chief AI officer Prashant Mehrotra tells American Banker the bank never ran token maxing. Every proposed use case is bucketed by whether it drives growth, cuts cost, or merely feels productive — and the last bucket loses.

tokenmaxxingcost-governanceai-spend
Read note