news

Introducing Claude Opus 5

Anthropic shipped Claude Opus 5 on July 24 priced level with Opus 4.8, pitching near-Fable 5 intelligence at half the price. Its own Frontier-Bench v0.1 figures show more than double the predecessor's score, and cheaper per task.

Published 2026-07-24Source: Anthropic
Anthropic source artwork

Why it matters

Vendor-reported numbers, but the framing is cost per task rather than raw score, which is what finance teams actually budget against. On OSWorld 2.0, Anthropic claims Opus 5 tops Fable 5's best mark while spending a shade over a third as much.

Tokenmaxxing read

The effort setting is the routing story inside one model: dial up intelligence or conserve tokens per call. Anthropic reports roughly 1.5x the next-best pass rate on Zapier's AutomationBench for the same cost per task, and more tasks passed than any rival even at lowest effort.

Source takeaway

Customer notes point the same way: a legal team reports matching quality on 26% fewer tokens than Opus 4.8 at max reasoning, while a trading firm says it burns roughly one-seventh as many reasoning tokens at under half the latency.

Topic links

tokenmaxxingcoding-agentstopicagentstopicscoreboards
Related projects

Tools that match this angle

#4In spirit
Agents

LangGraph

langchain-ai/langgraph

A framework for building resilient stateful agents with explicit graphs, persistence, human-in-the-loop flows, and controllable execution.

38.5K6.5KMIT
agentsstateworkflows
#5Direct
Evaluation

promptfoo

promptfoo/promptfoo

A CLI and CI workflow for testing prompts, agents, and RAG systems across models, with evals and red-team style checks.

23.8K2.1KMIT
prompt-evalscirag
#6In spirit
Evaluation

DSPy

stanfordnlp/dspy

A framework for programming and optimizing language-model pipelines rather than hand-tuning one prompt at a time.

36.5K3.1KMIT
optimizationprogrammingevals
Related feed

More source-linked context

Tech Times source artwork
newsTT
news

rtk Raises Claude Code Costs at Low Effort: JetBrains Benchmark Debunks 60–90% Claim

A JetBrains benchmark (July 20) ran ‘rtk,’ a proxy marketed to cut Claude Code tokens 60–90%, across 425 billed trials. At low effort it made sessions a median 7.6% MORE expensive—while rtk’s own analytics logged 96.2M tokens ‘saved.’

tokenmaxxingcoding-agentsagents
Read note
MakeUseOf source artwork
newsM
news

I taught Claude to talk like a "caveman" and ended up saving 70% of my output tokens

MakeUseOf benchmarked a community caveman plugin for Claude against normal and terse prompting, counting input and output tokens for a single prompt and then averaged across ten prompts.

tokenmaxxingcoding-agentsagents
Read note
Generated Tokenmaxxing editorial thumbnail for After ‘Tokenmaxxing’, Token Spend Has Become The New Metric To Watch
newsF
news

After ‘Tokenmaxxing’, Token Spend Has Become The New Metric To Watch

Forbes contributor Tim Keary argues the tokenmaxxing push has flipped into cost discipline: with CFOs and boards watching, firms now track token spend per engineer and blend premium and cheaper models instead of maximizing usage.

tokenmaxxingcoding-agentsagents
Read note