Authors: Barun Singh Bisht, Gurpreet Kaur, Gurpreet Kohli
Abstract: Automatic Speech Recognition (ASR) systems achieved significant progress in converting speech into text; however, most existing approaches focused primarily on lexical accuracy and overlooked the emotional context present in human speech. This limitation reduced the effectiveness of ASR in applications that required natural and expressive interaction. In this study, a pipeline was developed for preserving emotional information from speech through prosodic feature analysis. Raw audio was processed through extraction, enhancement, alignment, and feature analysis stages to generate a structured dataset with synchronized annotations. Two models, a Support Vector Machine (SVM) and a Bidirectional Long Short-Term Memory (BiLSTM) network, were trained on six emotion classes, namely happy, sad, neutral, angry, fear, and curious, to evaluate their performance in emotion recognition. The experimental results showed that while the SVM provided a reasonable baseline, reaching 50 percent accuracy at 1000 samples per emotion, the BiLSTM model achieved higher accuracy of 69 percent under the same conditions, owing to its ability to capture temporal dependencies in speech. These findings highlight the importance of prosodic features and sequential modelling for developing more expressive and context-aware speech recognition systems.
International Journal of Science, Engineering and Technology