2 papers
cs.CL2026
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi +2
In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 ho…
cs.CL2024
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
Ashwin Sankar, Srija Anand, Praveen Srinivasa Varadhan +7
Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is sca…