Authors: Jonnalagadda Somaiah, Professor Dr. Koppula Chinabusi
Abstract: Knee osteoporosis is a progressive bone disorder that increases fracture risk and significantly reduces mobility, making early diagnosis essential for effective treatment. Conventional deep learning approaches based solely on Convolutional Neural Networks (CNNs) often struggle to capture complex clinical patterns and contextual information present in medical data. This study proposes an explainable hybrid deep learning framework that integrates CNN-based feature extraction with fine-tuned Vision-Language Models (VLMs) to improve the automated classification of knee osteoporosis from X-ray images. Multiple CNN architectures, including ResNet50, DenseNet121, MobileNetV2, InceptionV3, Xception, and NASNetMobile, are evaluated and compared with a multimodal transformer-based model combining Vision Transformer (ViT) and BERT. The proposed framework incorporates transfer learning, multimodal feature fusion, hyperparameter optimization, and Grad-CAM-based visual interpretation to enhance diagnostic accuracy and model transparency. Experimental evaluation demonstrates that the fine-tuned multimodal model achieves superior classification performance with an accuracy of 93%, outperforming conventional CNN models while providing clinically interpretable predictions. The proposed framework offers a reliable, scalable, and explainable computer-aided diagnosis system that can support healthcare professionals in the early detection and severity assessment of knee osteoporosis.
International Journal of Science, Engineering and Technology