BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pydata-amsterdam2026//speaker//9VYCLY
BEGIN:VTIMEZONE
TZID:Europe/Amsterdam
BEGIN:DAYLIGHT
DTSTART:20250910T000000
TZNAME:CEST
TZOFFSETFROM:+0200
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20251026T030000
RDATE:20261025T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260329T030000
RDATE:20270328T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Reliable\, rigorous\, wrong: A psychometric view of LLM evals - Jo
 die Burchell
DTSTART;TZID=Europe/Amsterdam:20260910T115500
DTEND;TZID=Europe/Amsterdam:20260910T122500
DTSTAMP:20260911T153631Z
UID:pretalx-pydata-amsterdam2026-DK7SZW@pretalx.com
DESCRIPTION:Have you ever shipped a change because your eval score went fr
 om 76% to 85%\, only to wonder afterwards what that number actually proved
 ? Our evals are more sophisticated than ever: held-out test sets\, calibra
 ted LLM judges\, statistically significant differences\, and CI pipelines 
 running automatically on every commit. So we can trust that the score real
 ly did change. But can we trust the conclusion we're drawing from it: that
  the system got better?\n\n\nIn this talk\, we'll see that LLM evaluation 
 is not just an engineering problem\, but a measurement problem\, and one t
 hat psychology has been grappling with for over a century. Psychometrics h
 as spent decades asking how to measure things we can't observe directly\, 
 with the core concept of validity asking not just "is this score reliable?
 " but also "what evidence supports the interpretation I want to make from 
 it?" We’ll start by seeing how benchmarks fail basic validity checks\, i
 ncluding exploring this ourselves in a live demo. We’ll then bring the s
 ame perspective to the evals we build for real LLM applications\, where me
 trics such as “helpfulness”\, “quality” or “task success” can 
 be just as difficult to interpret.\n\n\nWe'll close with four validity que
 stions you can apply to your next eval\, giving you a concrete checklist f
 or thinking more carefully about what your metrics actually tell you. You'
 ll leave with a practical way to distinguish between treating a score as a
  fact about your application and treating it as evidence for a claim about
  it\, helping you build evals that tell you whether your changes are actua
 lly making the system better.
LOCATION:Entropy
URL:https://pretalx.com/pydata-amsterdam2026/talk/DK7SZW/
END:VEVENT
END:VCALENDAR
