long-form

20VC x SaaStr This Week : Apple Sues OpenAI, the Token-Maxing Era Begins, and the TAM Question Hanging Over AI Coding

On the 20VC/SaaStr podcast, Jason Lemkin, Harry Stebbings and Rory O'Driscoll call the budget era: ClickHouse has grown its AI bill 60-fold since February, while the best engineers keep 10 to 20 agents busy overnight.

Published 2026-07-16Source: SaaStrAI
SaaStrAI source artwork

Why it matters

Their prediction is that every company ends up splitting its stack in two — a budget model for routine work, a frontier one reserved for genuinely hard problems — the moment unlimited spend turns into an allowance. Meta aimed Spark 1.1 at that lower tier.

Tokenmaxxing read

The cheap tier is not a discount bin, it is the permanent high-volume slot — and it can be upgraded overnight by pouring more frontier reasoning into it, which is how Lemkin frames Haiku at a tenth of a cent. The ceiling math is the bear case sitting underneath.

Source takeaway

O'Driscoll's unresolved question: with only 1.8 million US developers earning roughly $250B a year in aggregate, labs whose enterprise revenue is mostly coding may already command a fifth of all US software wages.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

56.5K10.7KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

33.2K3.6KSource-available
tracesevalscosts
#4In spirit
Agents

LangGraph

langchain-ai/langgraph

A framework for building resilient stateful agents with explicit graphs, persistence, human-in-the-loop flows, and controllable execution.

39.9K6.7KMIT
agentsstateworkflows
Related feed

More source-linked context

Generated Tokenmaxxing editorial thumbnail for Bunq adopts Orq.ai router amid Europe AI sovereignty push - IT Brief UK
newsIB
news

Bunq adopts Orq.ai router amid Europe AI sovereignty push - IT Brief UK

IT Brief UK reports bunq replaced in-house LLM routing with Orq.ai’s router, citing rising maintenance costs and gaps in observability, governance, and performance.

tokenmaxxingcost-governanceai-spend
Read note
XDA source artwork
long-formX
long-form

Dropping Claude Code from High to Medium effort cut output tokens 45%

XDA's Mahnoor Faisal ran five coding jobs on Sonnet 5 twice from an identical starting codebase, changing only the effort level. High spent about 26,000 output tokens; Medium finished the same work on roughly 14,300.

coding-agentstoken-consumptionai-spend
Read note
Futurum source artwork
newsF
news

The End of Token Maxing: Why Pragmatic AI Engineering is Replacing Frontier Models

On Utilizing AI Ep. 37, Futurum analysts Brad Shimmin and Guy Currier argue enterprises are retiring default frontier models for smaller, quantized, task-specific ones placed behind abstraction layers and deterministic routers.

tokenmaxxingmodel-routingai-spend
Read note