11 papers
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
Pavan Kumar J, Agneedh Basu, Pranav Bhat +3
Most public sound-event corpora are optimized either for general audio tagging or for clean speech separation, and comparatively few provide strong timestamped noise annotations la…
Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages
Varun Rai, Pavan Kumar J, Sujith Pulikodan +1
Transformer-based models have shown strong accuracy in distinguishing spontaneous from scripted speech and natural from synthetic speech, but these results are established on a nar…
SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages
Sujith Pulikodan, Agneedh Basu, Pavan Kumar J +4
India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction o…
Audio--Image Alignment as a Continued-Pretraining Stage Improves Low-Resource ASR
Sujith Pulikodan, Nihar Desai, Prasanta Kumar Ghosh
Thousands of languages are spoken worldwide, yet many remain under-resourced for Automatic Speech Recognition (ASR) due to the limited availability of high-quality transcribed spee…
Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi
Sujith Pulikodan, Agneedh Basu, Saurabh Kumar +5
Benchmarking is critical for the systematic evaluation and comparison of automatic speech recognition (ASR) systems. While several open-source datasets are available for Hindi ASR,…
Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages
Pavan Kumar J, Agneedh Basu, Pranav Bhat +4
Self-supervised speech encoders are often fine-tuned with language supervision, which can overlook geographical variation. To understand the learned representations under joint sup…