collaborators

6 papers

cs.CL2025

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

Praveen Srinivasa Varadhan, Srija Anand, Soma Siddhartha +1

What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice…

cs.CL2025

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations

Ashwin Sankar, Yoach Lacombe, Sherry Thomas +3

We introduce RASMALAI, a large-scale speech dataset with rich text descriptions, designed to advance controllable and expressive text-to-speech (TTS) synthesis for 23 Indian langua…

cs.CL2024

Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

Praveen Srinivasa Varadhan, Amogh Gulati, Ashwin Sankar +8

Despite rapid advancements in TTS models, a consistent and robust human evaluation framework is still lacking. For example, MOS tests fail to differentiate between similar models,…

cs.CL2024

ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams

Srija Anand, Praveen Srinivasa Varadhan, Mehak Singal +1

Recent advancements in Text-to-Speech (TTS) technology have led to natural-sounding speech for English, primarily due to the availability of large-scale, high-quality web data. How…

cs.CL2024

IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS

Ashwin Sankar, Srija Anand, Praveen Srinivasa Varadhan +7

Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is sca…

cs.CL2024

Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings

Praveen Srinivasa Varadhan, Ashwin Sankar, Giri Raju +1

We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the…