
Building a Model Evaluation Practice for Production AI
About this event
Most teams pick an AI model the same way they pick a coffee machine - the demo looks good or they hear about it from someone, so they assume it works. But in production, evaluation decisions directly impact cost, latency, accuracy, and user trust. Without a structured eval practice, you're flying blind when a model regresses, or a new release claims to be better. This session gives you a framework for evaluating LLMs against your workloads and operationalizing evaluation as a continuous practice, not a one-time exercise. What we'll cover: • Why standard benchmarks are often misleading for enterprise use cases • Building a golden dataset: how to sample, label, and version evaluation sets from real production traffic • Key metrics beyond accuracy: hallucination rates, precision/recall tradeoffs, cost-per-correct-answer, and latency percentiles • Running evals at scale: open-source tools (Ragas, PromptFoo, LangSmith) and using Amazon Bedrock Model Evaluation • Setting a model evaluation gate: how to enforce a promotion threshold before any model goes to production • Demo: running a side-by-side eval across two models on a real classification task High level schedule: • 4:30 - 4:45: Arrivals • 4:45 - 5:45: Presentation & Demo • 5:45 - 6:00: Q&A • 6:00 - 6:30: Networking Important instructions • Event starts at 4:45 pm EST, but please allow at least 15 minutes for security to process registration • Please ensure that your meetup profile has your full name. Both first and last name are required and we will not be able to register attendees with just abbreviations or incomplete names. • There is an optional networking and Q&A event at the end of the meetup
Questions & comments
Ask the host anything — replies are visible to everyone.
—