Friday, October 9, 2026
5:30 PM CDT
Gavel & Gown: LLM Judge
About this event
An Organized AI session at Antler VC on the LLM-as-judge pattern — using language models to evaluate and score AI outputs — with practical guidance for builders and teams.
Evaluating AI with AI
The third installment of Organized AI's how-to series tackles a problem most teams hit once agents reach production: "looks good" is a vibe, not a QA process. Without a repeatable way to score output, weak answers ship and strong ones get discarded, and nobody can explain either call.
The session is a laptop-open build of an LLM-as-judge — a language model that evaluates another system's output against a rubric and returns a keep-or-revert verdict. Topics covered include making the process deterministic so identical input always yields the same ruling, scoring against an explicit rubric rather than intuition, and feeding verdicts back so the evaluation loop improves over time.
Aimed at founders, engineers, and AI operators building agents that need to be correct rather than merely plausible. Bring a laptop.
Attending
Held at Antler VC, 800 Brazos St #340, in downtown Austin, with recordings offered alongside the in-person session. Approval is required — requests are reviewed, so ask for a spot early rather than at the door.
One detail to confirm: the organizer's listing shows a start of 5:30 p.m. on October 9 but an end timestamp two days later, which does not match the single-evening format of the rest of this series. The start time is reliable; check the event page for the on-site finish time before planning around it.
Organized by
An Austin community hosting hands-on sessions and demos on applied AI workflows and tooling, from marketing automation to LLM-driven products, for builders and operators.