long-form

Dropping Claude Code from High to Medium effort cut output tokens 45%

XDA's Mahnoor Faisal ran five coding jobs on Sonnet 5 twice from an identical starting codebase, changing only the effort level. High spent about 26,000 output tokens; Medium finished the same work on roughly 14,300.

Published 2026-08-08Source: XDA
XDA source artwork

Why it matters

This is the cheapest lever available on an agent bill: one setting, no prompt rewriting and no model swap. Medium still completed all five jobs, so the saving did not quietly come out of task success.

Tokenmaxxing read

The savings scale with job size, which is the routing argument in miniature. The bug fix saved just 25% (2,000 tokens down to 1,500) while the search feature saved 48% (14,400 down to 7,500). Start at Medium and buy more reasoning when a task earns it.

Source takeaway

One developer, one project, five tasks — directional, not a benchmark. Worth rerunning on your own repo: both effort levels independently picked the same under-tested function, found the same bug, and shipped the same one-line fix with 12 tests green.

Topic links

Related projects

Tools that match this angle

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

59.3K11.6KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

34.9K3.8KSource-available
tracesevalscosts
#4In spirit
Agents

LangGraph

langchain-ai/langgraph

A framework for building resilient stateful agents with explicit graphs, persistence, human-in-the-loop flows, and controllable execution.

42.1K7.1KMIT
agentsstateworkflows
Related feed

More source-linked context

XDA source artwork
long-formX
long-form

Claude Code was using 51,000 tokens before I even typed a prompt — I fixed it

Mahnoor Faisal opened a new Claude Code session, ran /context, and found 51,400 tokens already loaded. Disabling four test plugins and auto-memory got the starting context down to roughly 41,400 before any real prompt.

coding-agentstoken-consumptiontoken-waste
Read note
Generated Tokenmaxxing editorial thumbnail for Enterprise AI budgets break at the handoff to production
long-formEC
long-form

Enterprise AI budgets break at the handoff to production

Express Computer interviews New Relic India's Ganesh Narasimhadevara on why AI bills keep climbing while the blended cost per million tokens has fallen over a year, from $18.40 in Q1 2025 down to $6.07 in Q1 2026.

tokenmaxxingai-spendcost-governance
Read note
Cisco Newsroom source artwork
newsCN
news

Cisco's Splunk adds Tokenomics to track coding-agent token spend

At Splunk .conf on Sept. 15, Cisco added a Tokenomics module to Splunk Agent Observability. It attributes token spend across AI agents and across employees' use of coding agents, naming Claude Code, Codex and Cursor.

ai-spendcoding-agentsllm-observability
Read note