long-form

20VC x SaaStr This Week : Apple Sues OpenAI, the Token-Maxing Era Begins, and the TAM Question Hanging Over AI Coding

On the 20VC/SaaStr podcast, Jason Lemkin, Harry Stebbings and Rory O'Driscoll call the budget era: ClickHouse has grown its AI bill 60-fold since February, while the best engineers keep 10 to 20 agents busy overnight.

Published 2026-07-16Source: SaaStrAI
SaaStrAI source artwork

Why it matters

Their prediction is that every company ends up splitting its stack in two — a budget model for routine work, a frontier one reserved for genuinely hard problems — the moment unlimited spend turns into an allowance. Meta aimed Spark 1.1 at that lower tier.

Tokenmaxxing read

The cheap tier is not a discount bin, it is the permanent high-volume slot — and it can be upgraded overnight by pouring more frontier reasoning into it, which is how Lemkin frames Haiku at a tenth of a cent. The ceiling math is the bear case sitting underneath.

Source takeaway

O'Driscoll's unresolved question: with only 1.8 million US developers earning roughly $250B a year in aggregate, labs whose enterprise revenue is mostly coding may already command a fifth of all US software wages.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

59.3K11.6KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

34.9K3.8KSource-available
tracesevalscosts
#4In spirit
Agents

LangGraph

langchain-ai/langgraph

A framework for building resilient stateful agents with explicit graphs, persistence, human-in-the-loop flows, and controllable execution.

42.1K7.1KMIT
agentsstateworkflows
Related feed

More source-linked context

Generated Tokenmaxxing editorial thumbnail for Enterprise AI budgets break at the handoff to production
long-formEC
long-form

Enterprise AI budgets break at the handoff to production

Express Computer interviews New Relic India's Ganesh Narasimhadevara on why AI bills keep climbing while the blended cost per million tokens has fallen over a year, from $18.40 in Q1 2025 down to $6.07 in Q1 2026.

tokenmaxxingai-spendcost-governance
Read note
Generated Tokenmaxxing editorial thumbnail for Bunq adopts Orq.ai router amid Europe AI sovereignty push - IT Brief UK
newsIB
news

Bunq adopts Orq.ai router amid Europe AI sovereignty push - IT Brief UK

IT Brief UK reports bunq replaced in-house LLM routing with Orq.ai’s router, citing rising maintenance costs and gaps in observability, governance, and performance.

tokenmaxxingcost-governanceai-spend
Read note
Cisco Newsroom source artwork
newsCN
news

Cisco's Splunk adds Tokenomics to track coding-agent token spend

At Splunk .conf on Sept. 15, Cisco added a Tokenomics module to Splunk Agent Observability. It attributes token spend across AI agents and across employees' use of coding agents, naming Claude Code, Codex and Cursor.

ai-spendcoding-agentsllm-observability
Read note