Authors: Dolly Chauhan, Kritika Sharma
Abstract: This paper presents the development of a real-time speech-to-speech translator and its analysis in order to overcome the language barriers between two languages. Based on advanced deep learning models, i.e. the Whisper-Speech-to-Text (STT) and MMS-Text-to-Speech (TTS) APIs of OpenAI and Facebook, respectively, the system takes the audio input, identifies the language spoken, converts it to English text and thereby transforms this text to English speech. The proposed architecture will take into account powerful audio preprocessing, such as Voice Activity Detection (VAD) and noise elimination, to improve the accuracy and reliability of the translation pipeline. A dataset of 200 samples from the Mozilla Common Voice corpus was the one used for the system's validation, which clearly showed its 16.6% latency decrease and superior noise robustness over the standard cascaded API baselines. The accuracy of translation, system time, and the quality of synthesized speech performance metrics are tested in a variety of languages, which proves the efficiency of the system and the possibility of its advancement into the practical side of using the system in various communication settings. The prototyped system presents the effectiveness of composing AI models in providing a smooth on-the-fly language conversion.
International Journal of Science, Engineering and Technology