Explainable Modality-Specific DenseNet Models and Population-Level Decision Fusion for Breast Cancer Classification Across Mammography, Ultrasound, and Histopathology

2 Sep

Authors: Muhammad Hamza, Sufiyan Hamza

Abstract: Breast cancer diagnosis draws on complementary information from mammography, ultrasound, and histopathology, yet multimodal artificial-intelligence studies can become clinically ambiguous when unrelated public cohorts are fused as though they represented the same patients. This study developed an explainable framework based on independently trained DenseNet-121 specialists for the three imaging modalities and evaluated feature- and decision-level fusion as population-level ensemble experiments. The harmonized analytical collection contained 34,207 publicly available images: 24,576 mammograms, 1,722 ultrasound images, and 7,909 histopathology images. Modality-specific preprocessing, class weighting, augmentation, and ImageNet transfer learning were applied, while the histopathology branch additionally incorporated a domain-adversarial objective. Grad-CAM++ was used for qualitative auditing of the radiological specialists. From the detailed test confusion matrices, ultrasound achieved 90.8% accuracy, 88.2% sensitivity, and 92.9% specificity; mammography achieved 96.0% accuracy, 96.3% sensitivity, and 95.6% specificity; and histopathology achieved 94.8% accuracy, 96.5% sensitivity, and 91.5% specificity. Wilson 95% confidence intervals, balanced accuracy, and Matthews correlation coefficients were derived from these counts. The domain-adversarial histopathology ablation showed an approximately two-percentage-point increase in rounded accuracy, although the archived experimental record did not preserve the precise source-target domain assignment. Attention-based feature fusion achieved 94.2% accuracy and assigned the largest learned weights to histopathology and mammography. Across eight exploratory fixed decision-weight configurations, the highest observed population-level accuracy was 98.8% for a mammography-emphasized rule. Because modality predictions originated from unmatched cohorts and the fixed-weight sweep was not documented as using a separate selection set, fusion findings are interpreted as exploratory sensitivity analyses rather than same-patient clinical performance. The results support transparent modality-specific evaluation and motivate future validation on deduplicated, patient-linked, externally collected cohorts with calibrated probabilities and reproducible split manifests.

DOI: https://doi.org/10.5281/zenodo.22246661