Why it matters
Benchmark scores drive model-routing and budget decisions; if the benchmarks are gamed or misleading, teams overspend on models that look strong on paper but underperform in production.
The tokenmaxxing angle
Talk covers eval failure modes (contamination, judge bias, agent scaffolding) that distort cost-per-quality comparisons across models — directly relevant to picking the cheapest model that actually clears your bar.
From the organizers
Speaker Rakshak Talwar names MMLU, GPQA, SWE-bench, GAIA, and OSWorld as the benchmarks under scrutiny; talk held in STATION Austin's Voltron Room, sponsored by STATION Austin.