Why it matters
Voice agents run on always-on streaming STT, so pipeline choices directly drive per-minute cost and latency; a live vendor benchmark plus real production builders comparing notes is a rare chance to see those tradeoffs argued in public.
The tokenmaxxing angle
AssemblyAI claims 10.2% lower word-error-rate across a 20,000-file benchmark for its new realtime model, plus built-in conversation memory -- exactly the per-call accuracy/latency/cost tradeoff that decides if a voice-agent stack is affordable at scale.
From the organizers
Panelists were Zhongren Shao (Sr. Software Engineer, Retell) and Adam Schuld (CTO, Super), moderated by Ryan Seams (VP, AssemblyAI); the demo covered AssemblyAI's Universal-3.5 Pro Realtime model with a stated 10.2% WER gain on a 20,000-file benchmark.