Authors: Richa Sharma, Archita Kumari, Archana Mongia
Abstract: Breast cancer diagnostic machine learning models need to be not only accurate but also interpretable to be deployed clinically with any reasonable degree of confidence. Although Random Forest, XGBoost and Neural Networks generate accurate predictions, their respective "black box" nature of operation can undermine the clinical community's willingness to trust the output. Explainable AI supplement methods such as SHAP and LIME have been developed in part to address this issue. Herein, we compare the SHAP & LIME measured explainability consistency of three different classifying algorithms (Random Forest, XGBoost, and Neural Network) as they were trained on the Breast Cancer Wisconsin dataset. Each algorithm's accuracy, F1, and Receiver Operating Characteristic Area Under Curve (ROC-AUC) were calculated to evaluate model classifying performance. Spearman rank correlation was calculated to evaluate explainability consistency as assessed by SHAP & LIME. Neural Networks exhibited the highest classifying accuracy (96.49%) and F1-score (0.9718); however, XGBoost provided the greatest consistency between SHAP & LIME (ρ = 0.9534), followed by Neural Network (ρ = 0.9049) and Random Forest (ρ = 0.8889). XGBoost provides the best overall balance of predictive accuracy with explainability consistency.
International Journal of Science, Engineering and Technology