Why it matters
Directly tackles the gap most teams hit after a demo works, namely how to know an agent is still behaving correctly once real users and real data hit it in production.
The tokenmaxxing angle
Covers instrumentation and LLM-as-judge evaluation for deployed agents, the observability layer that makes agent spend and quality regressions visible instead of guessed at.
From the organizers
The published agenda allots a dedicated 30-minute block to evaluation strategies and drift detection using LLM-as-judge, presented by AWS Solutions Architects Senthil Kamala Rathinam and Hasan Mehdi Rizvi.