Events / Austin

Benchmarks Are Lying to You: How to Evaluate Models and Agents in the Real World

Austin Deep Learning journal-club talk by Rakshak Talwar on evaluating LLMs and AI agents, examining benchmarks like MMLU, GPQA, SWE-bench, GAIA, and OSWorld and where they mislead.

Wed, Aug 5, 12:00 AMSTATION Austin, Voltron Room (1st floor) · Austin · TX

Why it matters

Benchmark scores drive model-routing and budget decisions; if the benchmarks are gamed or misleading, teams overspend on models that look strong on paper but underperform in production.

The tokenmaxxing angle

Talk covers eval failure modes (contamination, judge bias, agent scaffolding) that distort cost-per-quality comparisons across models — directly relevant to picking the cheapest model that actually clears your bar.

From the organizers

Speaker Rakshak Talwar names MMLU, GPQA, SWE-bench, GAIA, and OSWorld as the benchmarks under scrutiny; talk held in STATION Austin's Voltron Room, sponsored by STATION Austin.