Riccardo Cappuzzo
I am a research engineer at Inria, part of P16 and of the SODA research team. I am one of the maintainers of the skrub Python package. I hold a PhD in Computer Science and I am also interested in research on tabular learning and tabular foundational models.
Session
Skrub is a package that eases preparing dataframes so they can be used in machine-learning tasks. In practice, data can be spread over multiple tables, represent various types of information (numbers, text, dates, images...), or be stored on external database systems rather than dataframes. Crucially, data preparation often involves different steps at training or prediction time as stateful pre-processing steps (where part of the preparation is learned or adjusted on the train data e.g., imputation).
Skrub Data Ops help constructing complex pipelines to handle this variety of scenarios, while at the same time avoiding data leakage and allowing to tune the parameters of the full preprocessing pipeline to maximize the performance of the final machine learning model.
In this talk, we give a brief introduction of the Data Ops framework before presenting practical use cases, ranging from churn prediction, to energy usage forecasting, to image processing. Data Ops simplify greatly training models in advanced, multi-table situations.
By the end of the talk, attendees will learn about the skrub Data Ops to build data preparation and feature engineering pipelines that assemble rich transformations across multiple tables.