Francesco Bruzzesi
Data science tech lead at intella.tech · Mathematician at heart · Open source enthusiast
Session
If you do statistical work in Polars today, you leave Polars to do it: .to_numpy(), a trip through scipy.stats, and back. That cuts the lazy query in half and materializes everything. The other way out, a map_elements UDF, is slower still and holds the GIL.
scipy broadcasts parameter arrays fine, so a different distribution per row, where the mean and standard deviation are themselves columns, is already vectorized there. It hands the result back as a NumPy array, and the lazy plan ends there.
This talk walks through a Polars expression plugin that exposes scipy.stats-style distributions natively inside Polars expressions, with column-valued parameters as a first-class feature. The math is Rust (statrs); the surface is Polars. We will cover the plugin architecture (pyo3-polars), why column-valued parameters reshape the API, how seeded sampling stays reproducible across platforms, chunk layouts and both engines, and the null and error contract.
The thread running through it is what we got wrong before we got it right: the abstraction we built and then changed, the parameterization we flipped, the contract we had to rewrite. Expect concrete code, real benchmarks against scipy + polars.