Building an ML pipeline for churn prediction with skrub Data Ops
Skrub is a package that eases preparing dataframes so they can be used in machine-learning tasks. In practice, data can be spread over multiple tables, represent various types of information (numbers, text, dates, images...), or be stored on external database systems rather than dataframes. Crucially, data preparation often involves different steps at training or prediction time as stateful pre-processing steps (where part of the preparation is learned or adjusted on the train data e.g., imputation).
Skrub Data Ops help constructing complex pipelines to handle this variety of scenarios, while at the same time avoiding data leakage and allowing to tune the parameters of the full preprocessing pipeline to maximize the performance of the final machine learning model.
In this talk, we give a brief introduction of the Data Ops framework before presenting practical use cases, ranging from churn prediction, to energy usage forecasting, to image processing. Data Ops simplify greatly training models in advanced, multi-table situations.
By the end of the talk, attendees will learn about the skrub Data Ops to build data preparation and feature engineering pipelines that assemble rich transformations across multiple tables.
Building a machine learning pipeline is rarely a straightforward effort:
- Data can be spread across multiple sources and storage formats.
- Data preparation can involve multiple steps, unclear choices and assembling data coming from different sources.
- Operations may be stateful and require training, and their training must be performed only on the train set to avoid data leakage.
- Certain operations may need to be performed differently during training and prediction (e.g., fetching the required data may need be done differently)
- There may be hyper-parameters to tune.
- At the end of the process, it should be possible to re-execute all the same operations with the same parameters on unseen data.
Skrub Data Ops are a pipeline-building framework that alleviates these criticalities: Data Ops wrap around any arbitrary function provided by the user, including complex data fetching and preparation steps; they simplify combining tables by letting users adopt the dataframe library of their choosing; they store the state of any transformer in the pipeline and keep track of samples to avoid data leakage; they simplify the construction of complex hyper-parameter search spaces that allow setting arbitrary operations as choices.
This talk will present the Data Ops in multiple practical scenarios:
- Prediction of churn in a subscription-based online multiplayer game.
- Forecasting the electricity usage in France with weather and calendar features.
- Image classification with pytorch.
In each scenario we highlight how Data Ops deal with stateful transformations, hyperparmeter tuning, and the risk of data leakage.
All material for the talk will be made available online.
I am a research engineer at Inria, part of P16 and of the SODA research team. I am one of the maintainers of the skrub Python package. I hold a PhD in Computer Science and I am also interested in research on tabular learning and tabular foundational models.