Events / New York

Research Meetup: Benchmarking Coding Agents (NYC)

NYC lightning-talk meetup from Scale AI's Labs Research series on benchmarking coding agents, with four short talks introducing new evaluation benchmarks.

Wed, Aug 12, 5:30 PMWorld Trade Center, New York, NY

Why it matters

Benchmark design determines what “good” looks like for coding agents you might deploy or route between; understanding how these evals are built helps you judge vendor claims instead of trusting marketing.

The tokenmaxxing angle

New benchmarks like Terminal Bench and SWE Atlas are exactly what you'd use to justify a routing policy (cheap model for X, frontier model for Y) with data instead of guesswork; this is where those numbers get made.

From the organizers

Third in Scale's Labs Research Meetup Series; confirmed lightning talks: Kilian Lieret (Meta AI) on ProgramBench, Eugene Wu (Columbia) on BranchBench, Miguel Romero Calvo (Scale AI) on Terminal Bench, Mohit Raghavendra (Scale AI) on SWE Atlas.