Why it matters
Teams that skip evaluation and observability don't just ship worse agents — they burn budget on silent failures and reruns nobody is tracking down; reliability and cost are the same problem here.
The tokenmaxxing angle
The talk frames evaluation and observability as the layer deciding whether a multi-agent pipeline survives production — exactly the instrumentation gap that turns unexplained model spend into a mystery line on the API bill.
From the organizers
Speaker is Muazma Zahid, Group Product Manager at Google; session covers grounding techniques, evaluation frameworks, observability practices, and design patterns for reliable agent orchestration.