8 papers · 1 filter
Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages
Varun Rai, Pavan Kumar J, Sujith Pulikodan +1
Transformer-based models have shown strong accuracy in distinguishing spontaneous from scripted speech and natural from synthetic speech, but these results are established on a nar…
SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J +4
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction o…
Audio--Image Alignment as a Continued-Pretraining Stage Improves Low-Resource ASR
Sujith Pulikodan, Nihar Desai, Prasanta Kumar Ghosh
Thousands of languages are spoken worldwide, yet many remain under-resourced for Automatic Speech Recognition (ASR) due to the limited availability of high-quality transcribed spee…
Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi
Sujith Pulikodan, Agneedh Basu, Saurabh Kumar +5
Benchmarking is critical for the systematic evaluation and comparison of automatic speech recognition (ASR) systems. While several open-source datasets are available for Hindi ASR,…
Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages
Pavan Kumar J, Agneedh Basu, Pranav Bhat +4
Self-supervised speech encoders are often fine-tuned with language supervision, which can overlook geographical variation. To understand the learned representations under joint sup…
An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages
Sujith Pulikodan, Agneedh Basu, Pavan Kumar +4
Synthetic data has the potential to be a valuable resource for training machine learning models, particularly Automatic Speech Recognition (ASR) Systems; however, its effectiveness…