This event has ended. It stays here for the record. See upcoming Meetups → · Next from Austin Deep Learning
Aug
4

Tuesday, August 4, 2026
7:00 PM – 9:00 PM CDT

Benchmarks Are Lying to You: How to Evaluate Models and Agents in the Real World

MeetupFreeIntermediateDevelopersResearchers

About this event

An Austin Deep Learning talk on why standard AI benchmarks mislead — covering contamination, evaluator bias, and unfair comparisons — and how to build evaluations that actually predict production behavior.

An Austin Deep Learning technical session on the gap between benchmark scores and real-world performance. The talk surveys the major evaluation suites used across language models, coding systems, and autonomous agents — MMLU, GPQA, SWE-bench, GAIA, and OSWorld — and examines what each one does and does not actually measure.

From there it turns to the failure modes that make published numbers unreliable: training-data contamination, evaluator bias in model-graded scoring, and comparisons drawn between systems that were never evaluated under the same conditions. The second half addresses the practical question of building your own production-grade evaluations rather than trusting a leaderboard.

Attendees should come away better able to read benchmark results critically, compare models fairly, and validate AI systems against their real deployment conditions. Presented by Rakshak T. and Eric T. H. at STATION Austin, Voltron Room on the 1st floor, 701 Brazos Street.

Organized by

A
Austin Deep Learning

An Austin meetup community for machine learning practitioners and researchers, hosting technical talks on deep learning, language models, evaluation methodology, and agent systems.

All events by Austin Deep Learning

Visit organizer →