3 papers
cs.SD2023
RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations
Neha Sahipjohn, Neil Shah, Vishal Tambrahalli +1
Significant progress has been made in speaker dependent Lip-to-Speech synthesis, which aims to generate speech from silent videos of talking faces. Current state-of-the-art approac…
cs.SD2023
MParrotTTS: Multilingual Multi-speaker Text to Speech Synthesis in Low Resource Setting
Neil Shah, Vishal Tambrahalli, Saiteja Kosgi +2
We present MParrotTTS, a unified multilingual, multi-speaker text-to-speech (TTS) synthesis model that can produce high-quality speech. Benefiting from a modularized training parad…
cs.CL2023
ParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations
Neil Shah, Saiteja Kosgi, Vishal Tambrahalli +3
We present ParrotTTS, a modularized text-to-speech synthesis model leveraging disentangled self-supervised speech representations. It can train a multi-speaker variant effectively…