BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pydata-amsterdam2026//speaker//R9DMSA
BEGIN:VTIMEZONE
TZID:Europe/Amsterdam
BEGIN:DAYLIGHT
DTSTART:20250912T000000
TZNAME:CEST
TZOFFSETFROM:+0200
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20251026T030000
RDATE:20261025T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260329T030000
RDATE:20270328T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Beyond LLM-as-Judge: Using GLIDE for Reliable\, Scalable Evaluatio
 n of GenAI Systems - Grégoire Martinon
DTSTART;TZID=Europe/Amsterdam:20260912T100000
DTEND;TZID=Europe/Amsterdam:20260912T110000
DTSTAMP:20260809T115145Z
UID:pretalx-pydata-amsterdam2026-ZSYGEC@pretalx.com
DESCRIPTION:Agentic systems are booming\, yet few reach production. A prim
 ary blocker is performance evaluation on unstructured outputs (typically t
 ext): LLM-as-Judge is cheap but biased while human annotations don't scale
 . To bridge this gap\, practitioners need scalability without sacrificing 
 statistical rigor.\n\nThis talk introduces Prediction-Powered Inference\, 
 a framework that combines a small set of human labels with a large volume 
 of proxy judgments to produce unbiased performance measurements with forma
 l confidence intervals. We will showcase GLIDE\, a new Python open-source 
 package that makes state-of-the-art techniques like PPI++ accessible to ev
 ery developer. \n\nTarget Audience: ML and Software Engineers struggling t
 o evaluate GenAI systems\, and researchers in the field of prediction-powe
 red inference. No background required.\n\nTakeaways & Materials: All code 
 and tutorials\, will be shared via our open-source repository.
LOCATION:Room A
URL:https://pretalx.com/pydata-amsterdam2026/talk/ZSYGEC/
END:VEVENT
END:VCALENDAR
