Authors: Shivani Chauhan, Rimmy, Ashish Prajapati
Abstract: Speech-to-text (STT) and multi-speaker diarization have emerged as crucial components in intelligent communication systems, virtual meeting platforms, assistive technologies, and large-scale multimedia analytics. Recent advancements in transformer-based architectures such as Whisper and self-supervised pipelines like Pyannote have significantly improved transcription quality and speaker discrimination, enabling highly accurate multi-speaker processing even on consumer hardware. A system that integrates Whisper for multilingual transcription, Pyannote for diarization, and a new speaker identity recognition module that uses voice embedding is presented in this research as a modular, scalable, and real-time integrated system. Real-time microphone input, audio uploads, and live transcript visualization are all supported by a full-stack web platform that uses React and Flask. The Word Error Rate (WER) of 6.2% and the Diarization Error Rate (DER) of 13% were observed in experiments conducted on 5 hours of multi-speaker audio. 87% identity accuracy is achieved by personalized speaker recognition. Practical usability is demonstrated by the proposed system for meetings, lectures, podcasts, and automated captioning.
International Journal of Science, Engineering and Technology