Online Probability Recalibration for Machine Learning under Distribution Shift - A Study of Calibration Maintenance for In-hospital Mortality Risk Prediction
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Machine learning models used for intensive care unit (ICU) mortality prediction output probabilities that may guide monitoring and treatment decisions in a high stakes clinical setting. After deployment, however, distribution shift can make these probabilities miscalibrated, while online correction is constrained by delayed outcome labels and a limited rolling labelled buffer. This thesis studies online post-hoc recalibration under these conditions using the eICU Collaborative Research Database. A base predictor trained on 2014 data is kept fixed, while a separate calibration layer may update during deployment on a 2015 evaluation stream. Twelve update strategies are evaluated across three base models, four buffer sizes, and eleven shift regimes spanning the observed 2014–2015 shift, which was mild in this cohort, and ten controlled score-level drift scenarios. Under the observed mild shift, no trigger-based strategy improved on fixed calibration, and S12 Adaptive Hybrid remained inactive with zero refits across all three models. Under controlled drift, fixed calibration deteriorated, while trigger-based updating became beneficial. Across all eleven regimes, S12 achieved the lowest overall mean rank of 2.27, the narrowest spread of 0.92, and remained in the top three in every regime. Under drift, it used 15–31 refits per configuration, compared with 124 for the periodic baseline. These findings show that the value of online recalibration is regime-dependent: under mild shift, preserving the existing calibrator is sufficient, whereas under stronger or more structured drift, selective trigger-based updating becomes beneficial. Overall, the adaptive hybrid policy, which combines label-free and reliability-aware signals and adapts their balance to buffer maturity, provided the most consistent and efficient recalibration performance. These findings matter beyond this specific ICU mortality task because many deployed machine learning systems face similar practical constraints. Predictions are available immediately, but reliable outcome feedback arrives late, and retraining the full model maybe difficult, costly, or undesirable in regulated settings. In such cases, maintaining trustworthy probabilities is not only a matter of choosing a good calibrator, but also of deciding when an update is justified. The results support recalibration as a selective maintenance policy, offering a practical middle ground between doing nothing and updating continuously when the future drift regime is unknown.