Grégoire Martinon
With a career rooted in the rigors of General Relativity, I specialized in building robust, ethical, and scalable AI systems. I have several years of experiences contributing to the open-source community, having led the development of MAPIE, the reference library in conformal predictions. I now work on reliable and scalable evaluation of GenAI systems, and took the lead of GLIDE, a new open-source package implementing state-of-the-art semi-supervised inference techniques.
Session
Agentic systems are booming, yet few reach production. A primary blocker is performance evaluation: LLM-as-Judge is cheap but biased, while human annotation doesn't scale. The result is a familiar gap, automated metrics that look great on paper, and business experts who remain unconvinced.
This talk introduces GLIDE (Generated Label Inference and Debiasing Engine), an open-source Python package that bridges this gap using Prediction-Powered Inference. By combining a small set of human labels (~100 annotations) with large volumes of proxy judgments, GLIDE produces unbiased performance estimates with valid confidence intervals, powered by state-of-the-art techniques including Prediction-Powered Inference (PPI++), Active Statistical Inference, and a suite of samplers that optimize where your annotation budget is spent.
After this talk, you will know what statistically rigorous GenAI system evaluation looks like and how to put it in place
Target audience: ML and Software Engineers evaluating GenAI systems, and researchers in statistical inference. No prior background required.
All code, tutorials, and examples are available in the open-source repository.