Why it matters
Directly addresses how teams justify a model choice with evidence instead of vibes, which is the same evaluation discipline that should precede any model-routing or cost decision.
The tokenmaxxing angle
Talk covers building golden datasets from production traffic and tracking cost-per-correct-answer and latency percentiles as eval metrics -- a near-exact match for token-cost and model-routing decisions.
From the organizers
Description names specific eval tooling covered: Ragas, PromptFoo, LangSmith, and Amazon Bedrock Model Evaluation, plus a live side-by-side eval demo on a real classification task.