12 papers
SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J +4
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction o…
Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion
Saurabh Kumar, Amartyaveer, Prasanta Kumar Ghosh
Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences.…
Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings
Jesuraj Bandekar, Prasanta Kumar Ghosh
Acoustic-to-Articulatory Inversion (AAI) estimates vocal tract articulator movements from speech, benefiting tasks like ASR, speech synthesis, and speaker verification. While deep…
Audio--Image Alignment as a Continued-Pretraining Stage Improves Low-Resource ASR
Sujith Pulikodan, Nihar Desai, Prasanta Kumar Ghosh
Thousands of languages are spoken worldwide, yet many remain under-resourced for Automatic Speech Recognition (ASR) due to the limited availability of high-quality transcribed spee…
Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi
Sujith Pulikodan, Agneedh Basu, Saurabh Kumar +5
Benchmarking is critical for the systematic evaluation and comparison of automatic speech recognition (ASR) systems. While several open-source datasets are available for Hindi ASR,…
Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages
Pavan Kumar J, Agneedh Basu, Pranav Bhat +4
Self-supervised speech encoders are often fine-tuned with language supervision, which can overlook geographical variation. To understand the learned representations under joint sup…