Beyond the Holdout: Mitigating Censoring Bias with Asymmetric IPW
Large scale transactional platforms use machine learning models to block fraud before it occurs. However, once a model changes which transactions are allowed, it also changes which labels can be observed. The model cannot see outcomes of transactions it considered risky. Metrics from the surviving population become biased, and retraining on those labels can reinforce the same blind spots.
A randomized holdout restores an unbiased sample of the affected population, but creates a second problem: how should a small holdout be combined with the larger censored dataset? This talk compares six retraining strategies in a controlled simulation built with LightGBM. It shows when simply dropping censored labels is preferable, when inverse propensity weighting (IPW) recovers hidden information, and why neither method dominates across data regimes.
We then introduce Asymmetric IPW, which uses pooled holdout validation to choose between two retraining methods: dropping and pure IPW. Across six scenarios, it tracks the stronger method without requiring the underlying regime to be diagnosed. Attendees will learn why a holdout is necessary for unbiased evaluation, how censoring affects retraining, and how validation can balance bias and variance when retraining a model after launch.
Topic and Relevance
Censoring caused by a model policy occurs when its action prevents the predicted outcome from being observed. In fraud systems, blocked transactions never produce counterfactual fraud labels. Metrics from surviving transactions can then give the wrong answer, and retraining on those labels can reproduce the previous model's blind spots.
This controlled synthetic experiment preserves the true outcome before applying the policy. It provides an oracle for evaluation while recreating the labels available after deployment. Six scenarios vary global drift, hidden regional signal, regional drift, trigger rate, holdout size, and base rate. Every method uses the same LightGBM configuration and is evaluated with normalized partial AUC through 20% FPR.
Target Audience
The talk is intended for data scientists and machine learning engineers working in fraud, risk, credit, moderation, recommendations, or any setting where model actions affect which labels remain observable.
Audience Takeaways
- Understand how a policy can bias production metrics and future training data by censoring labels.
- Recognize why a randomized holdout is needed for unbiased evaluation after deployment.
- Understand the tradeoff between giving every uncensored label equal weight and using inverse propensity weighting on flagged holdout observations.
- Learn how Asymmetric IPW uses pooled holdout validation to choose between those methods without requiring the data regime to be diagnosed manually.
Talk Type and Approach
This is a conceptual and experimental talk. Visual examples introduce the censoring mechanism. The simulation then compares no retraining, incremental learning, no holdout, holdout only, dropping, and pure IPW. Asymmetric IPW selects between the two strongest methods.
Dropping and pure IPW use the same training rows. Dropping gives every row equal weight. Pure IPW upweights flagged holdout observations by the inverse holdout probability. This isolates the tradeoff between the lower variance of dropping and the lower bias of IPW.
The results are presented with paired comparisons and 95% confidence intervals. At the late endpoint, dropping clearly wins in three scenarios, IPW clearly wins in two, and one comparison is a statistical tie with a numerical IPW lead. Asymmetric IPW makes one choice from pooled holdout validation, without access to test performance. In retrospective evaluation, that choice matches the endpoint with the higher mean test score in all six late scenarios and statistically ties the fixed better endpoint in all twelve scenario and period comparisons. The same winner-matching pattern was observed in a real fraud dataset.
Required Background Knowledge
Basic familiarity with supervised classification and model evaluation is useful. LightGBM, inverse propensity weighting, and normalized partial AUC are mentioned in the talk.
Scope and Structure
- Censoring caused by model policies: 5 minutes
- Simulation design and assumptions: 4 minutes
- Why unbiased measurement requires a holdout: 5 minutes
- Retraining methods and the dropping versus IPW tradeoff: 6 minutes
- Asymmetric IPW, results, and limitations: 5 minutes
- Q&A: 5 minutes
The source files, generated results, and executed notebooks will be published on GitHub and shared during the presentation.
Senior Data Scientist based in Tallinn, Estonia, with 9+ years of experience in machine learning, analytics, experimentation, and causal inference.