Why it matters
The talks target the eval problems that determine whether agent spend is justified — knowing where in a multi-step task an agent wastes calls, not just scoring the final answer's accuracy.
The tokenmaxxing angle
Process-level error localization — pinpointing which step in a multi-agent chain went wrong rather than just the final output — is exactly the missing piece for catching where token-burning retries happen.
From the organizers
Hosted by MiniMax at Upscale in SF's Financial District, agenda runs 6-9pm with 10-15 minute research talks, a panel/Q&A, and is presented with partners Protégé and Together AI.