5 papers
Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
Kush Juvekar, Kavya Manohar, Aditya Srinivas Menon +2
Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio performance. To diagnose this mismatch, we int…
Scalable Offline ASR for Command-Style Dictation in Courtrooms
Kumarmanas Nethil, Vaibhav Mishra, Kriti Anandan +1
We propose an open-source framework for Command-style dictation that addresses the gap between resource-intensive Online systems and high-latency Batch processing. Our approach use…
Romanized to Native Malayalam Script Transliteration Using an Encoder-Decoder Framework
Bajiyo Baiju, Kavya Manohar, Leena G Pillai +1
In this work, we present the development of a reverse transliteration model to convert romanized Malayalam to native script using an encoder-decoder framework built with attention-…
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations
Kavya Manohar, Leena G Pillai, Elizabeth Sherly
This paper explores the pitfalls in evaluating multilingual automatic speech recognition (ASR) models, with a particular focus on Indic language scripts. We investigate the text no…
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
Leena G Pillai, Kavya Manohar, Basil K Raju +1
This paper presents a novel multistage fine-tuning strategy designed to enhance automatic speech recognition (ASR) performance in low-resource languages using OpenAI's Whisper mode…