6 papers
Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages
Chowdam Venkata Kumar, Kumud Tripathi, Pankaj Wasnik
Multilingual ASR models such as Whisper perform well on high-resource languages but exhibit substantially higher Word Error Rates (WER) for Dravidian languages compared to Indo-Ary…
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
Aditya Srinivas Menon, Kumud Tripathi, Raj Gohil +1
Self-supervised learning (SSL) has advanced speech processing but suffers from quadratic complexity due to self-attention. To address this, SummaryMixing (SM) has been proposed as…
Listen Like a Teacher: Mitigating Whisper Hallucinations using Adaptive Layer Attention and Knowledge Distillation
Kumud Tripathi, Aditya Srinivas Menon, Aman Gaurav +2
The Whisper model, an open-source automatic speech recognition system, is widely adopted for its strong performance across multilingual and zero-shot settings. However, it frequent…
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
Aditya Srinivas Menon, Raj Prakash Gohil, Kumud Tripathi +1
Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits s…
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
Kumud Tripathi, Chowdam Venkata Kumar, Pankaj Wasnik
Voice Activity Detection (VAD) plays a key role in speech processing, often utilizing hand-crafted or neural features. This study examines the effectiveness of Mel-Frequency Cepstr…
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
Kumud Tripathi, Raj Gothi, Pankaj Wasnik
Automatic speech recognition has recently seen a significant advancement with large foundational models such as Whisper. However, these models often struggle to perform well in low…