6 papers
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Adarsh Arigala, Arjun Gangwar, S Umesh +1
Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language understanding. Grounding text in its visual fo…
HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec
Arjun Gangwar, S Umesh
The popularity of neural audio codecs as speech tokenizers has surged with the advent of Multimodal Large Language Models. New codec architectures with semantic and acoustic disent…
Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations
Naman Kothari, Arjun Gangwar, Adarsh Arigala +1
Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual i…
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
Advait Joglekar, Divyanshu Singh, Rooshil Rohit Bhatia +1
Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectur…
Shiksha: A Technical Domain focused Translation Dataset and Model for Indian Languages
Advait Joglekar, Srinivasan Umesh
Neural Machine Translation (NMT) models are typically trained on datasets with limited exposure to Scientific, Technical and Educational domains. Translation models thus, in genera…
SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task
Hamees Sayed, Advait Joglekar, Srinivasan Umesh
We develop a robust translation model for four low-resource Indic languages: Khasi, Mizo, Manipuri, and Assamese. Our approach includes a comprehensive pipeline from data collectio…