SciPy 2026

Jackie Valeri

Hi! I am a Senior Data Scientist at Moderna working on machine learning and data science for pre-clinical research initiatives. I love custom algorithm development, iteratively designing libraries of molecules, and working with experimentalists to execute drug discovery campaigns. When I'm not in front of a screen, I love puzzling, baking, skiing, watching baseball, and reading sci fi/fantasy novels.


Session

07-16
13:15
30min
Just throw it away? Class imbalance lessons from molecular machine learning to meatballs
Jackie Valeri

Imbalanced datasets are common across science and industry: most screened molecules are inactive and most batted balls in baseball result in outs. One standard practice is to downsample the majority class or avoid collecting more of it. But majority-class examples are not interchangeable. Some are closely related to other examples, while others are distinct from any other example in the dataset. Others define the boundary between success and failure.

This talk asks two practical questions:

  1. How much majority-class data is actually necessary for a performative machine learning model?
  2. If we cannot collect all of it, which majority-class examples should we collect?

Using three wildly different datasets—antibacterial molecular screening, sandwich taste ratings, and Major League Baseball at-bat outcomes—I compare random downsampling to strategies that retain harder or more diverse majority-class examples, and evaluate the impact on generalization and performance for real-world machine learning models.

Data-Driven Discovery, Machine Learning and Artificial Intelligence
Memorial Hall