long-form

Enterprise AI budgets break at the handoff to production

Express Computer interviews New Relic India's Ganesh Narasimhadevara on why AI bills keep climbing while the blended cost per million tokens has fallen over a year, from $18.40 in Q1 2025 down to $6.07 in Q1 2026.

Published 2026-09-18Source: Express Computer
Generated Tokenmaxxing editorial thumbnail for Enterprise AI budgets break at the handoff to production

Why it matters

The FinOps Foundation's 2026 survey — 1,190-plus practitioners responsible for roughly $83 billion of annual cloud spend — found 73% of organisations overshot their AI cost projections. Just 43% have a formal AI governance policy at all.

Tokenmaxxing read

Routing is the cheapest fix on the table: tiered architectures report a median blended $2.31 per million tokens against $18.40 for default-to-frontier shops, and one enterprise cut a $40,000 monthly bill to $24,000 without changing the product.

Source takeaway

The hard data here is the rework. New Relic's survey of 200 US technology decision-makers found 78% saw production incidents spike and 82% hit a major failure tied to AI code within six months. Treat the unattributed spend anecdotes with more caution.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

59.1K11.6KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

34.8K3.8KSource-available
tracesevalscosts
#4In spirit
Agents

LangGraph

langchain-ai/langgraph

A framework for building resilient stateful agents with explicit graphs, persistence, human-in-the-loop flows, and controllable execution.

41.9K7.1KMIT
agentsstateworkflows
Related feed

More source-linked context

Cisco Newsroom source artwork
newsCN
news

Cisco's Splunk adds Tokenomics to track coding-agent token spend

At Splunk .conf on Sept. 15, Cisco added a Tokenomics module to Splunk Agent Observability. It attributes token spend across AI agents and across employees' use of coding agents, naming Claude Code, Codex and Cursor.

ai-spendcoding-agentsllm-observability
Read note
SaaStrAI source artwork
long-formS
long-form

20VC x SaaStr This Week : Apple Sues OpenAI, the Token-Maxing Era Begins, and the TAM Question Hanging Over AI Coding

On the 20VC/SaaStr podcast, Jason Lemkin, Harry Stebbings and Rory O'Driscoll call the budget era: ClickHouse has grown its AI bill 60-fold since February, while the best engineers keep 10 to 20 agents busy overnight.

tokenmaxxingai-spendmodel-routing
Read note
Generated Tokenmaxxing editorial thumbnail for AWS maps three layers of Bedrock token-cost visibility
guideAW
guide

AWS maps three layers of Bedrock token-cost visibility

Rachanee Singprasong and Chitresh Saxena lay out three layers of Bedrock cost visibility: native CloudWatch metrics plus CUR 2.0 caller-identity attribution, then model invocation logging, then OpenTelemetry in the client.

ai-spendllm-observabilitycost-governance
Read note