BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pyconde-pydata-2026//talk//3UHPZB
BEGIN:VTIMEZONE
TZID:Europe/Berlin
BEGIN:DAYLIGHT
DTSTART:20250414T000000
TZNAME:CEST
TZOFFSETFROM:+0200
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20251026T030000
RDATE:20261025T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260329T030000
RDATE:20270328T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:It Works on My Machine: Why LLM Apps Fail Users (Not Tests) - Thom
 as Prexl\, Frank Rust
DTSTART;TZID=Europe/Berlin:20260414T171000
DTEND;TZID=Europe/Berlin:20260414T174000
DTSTAMP:20260812T210837Z
UID:pretalx-pyconde-pydata-2026-3UHPZB@pretalx.com
DESCRIPTION:LLM applications frequently pass tests but fail users in produ
 ction. This talk examines the gap between evaluation metrics and user expe
 rience through three lenses: **Expectations** (what "working" means to use
 rs)\, **Functional** (system-level vs. component-level success)\, and **Op
 erational** (real-world reliability).\n\nDrawing from production experienc
 e\, we'll share scenarios of expectation mismatches\, silent failures\, an
 d undetected drift—plus practical strategies for bridging the gap. The c
 ore message: evaluation should answer whether your system serves users\, n
 ot whether it passes tests.
LOCATION:Palladium [2nd Floor]
URL:https://pretalx.com/pyconde-pydata-2026/talk/3UHPZB/
END:VEVENT
END:VCALENDAR
