Tuesday, August 4, 2026
7:00 PM – 9:00 PM CDT
Benchmarks Are Lying to You: How to Evaluate Models and Agents in the Real World
About this event
An Austin Deep Learning talk on why standard AI benchmarks mislead — covering contamination, evaluator bias, and unfair comparisons — and how to build evaluations that actually predict production behavior.
An Austin Deep Learning technical session on the gap between benchmark scores and real-world performance. The talk surveys the major evaluation suites used across language models, coding systems, and autonomous agents — MMLU, GPQA, SWE-bench, GAIA, and OSWorld — and examines what each one does and does not actually measure.
From there it turns to the failure modes that make published numbers unreliable: training-data contamination, evaluator bias in model-graded scoring, and comparisons drawn between systems that were never evaluated under the same conditions. The second half addresses the practical question of building your own production-grade evaluations rather than trusting a leaderboard.
Attendees should come away better able to read benchmark results critically, compare models fairly, and validate AI systems against their real deployment conditions. Presented by Rakshak T. and Eric T. H. at STATION Austin, Voltron Room on the 1st floor, 701 Brazos Street.
Organized by
An Austin meetup community for machine learning practitioners and researchers, hosting technical talks on deep learning, language models, evaluation methodology, and agent systems.