Why it matters
Vendor-reported numbers, but the framing is cost per task rather than raw score, which is what finance teams actually budget against. On OSWorld 2.0, Anthropic claims Opus 5 tops Fable 5's best mark while spending a shade over a third as much.
Tokenmaxxing read
The effort setting is the routing story inside one model: dial up intelligence or conserve tokens per call. Anthropic reports roughly 1.5x the next-best pass rate on Zapier's AutomationBench for the same cost per task, and more tasks passed than any rival even at lowest effort.
Source takeaway
Customer notes point the same way: a legal team reports matching quality on 26% fewer tokens than Opus 4.8 at max reasoning, while a trading firm says it burns roughly one-seventh as many reasoning tokens at under half the latency.


