Evaluating The Effects Of Missing-data Imputation Methods On Predictive Model Performance

5 Dec

Authors: Luke Akpan

Abstract: Background: Missing observations can alter fitted relationships, predictive accuracy and uncertainty, yet method comparisons often emphasize a single metric. Aim: To compare listwise deletion, mean/median, regression, k-nearest-neighbour (KNN) and multiple imputation (MI) across realistic prediction settings. Settings and design: A reproducible Monte Carlo simulation, parameterized to resemble a mixed clinical risk dataset, was conducted in May 2022. Materials and methods: Complete datasets were generated under linear and nonlinear signal structures; 10%, 30% or 50% values were removed under MCAR, MAR or MNAR mechanisms. Logistic regression and random forest models were evaluated in 1,000 repetitions. Statistical analysis: Area under the receiver-operating-characteristic curve (AUC), Brier score, calibration slope, root-mean-square error, confidence-interval coverage and computational failure were summarized. Results: MI was most reliable under MAR, retaining 96-99% of complete-data AUC and near-nominal coverage. KNN was competitive for nonlinear random forests at 10-30% missingness. Single regression imputation preserved discrimination but produced overconfident inference. Listwise deletion deteriorated rapidly as missingness increased; mean/median imputation showed systematic calibration loss. No method fully corrected MNAR bias. Conclusion: Imputation should be selected according to the prediction objective, missingness mechanism, model geometry and uncertainty requirement; sensitivity analysis is essential when MNAR is plausible.

DOI: http://doi.org/10.5281/zenodo.21735816