Why it matters
Benchmark design determines what “good” looks like for coding agents you might deploy or route between; understanding how these evals are built helps you judge vendor claims instead of trusting marketing.
The tokenmaxxing angle
New benchmarks like Terminal Bench and SWE Atlas are exactly what you'd use to justify a routing policy (cheap model for X, frontier model for Y) with data instead of guesswork; this is where those numbers get made.
From the organizers
Third in Scale's Labs Research Meetup Series; confirmed lightning talks: Kilian Lieret (Meta AI) on ProgramBench, Eugene Wu (Columbia) on BranchBench, Miguel Romero Calvo (Scale AI) on Terminal Bench, Mohit Raghavendra (Scale AI) on SWE Atlas.