Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Audio--Image Alignment as a Continued-Pretraining Stage Improves Low-Resource ASR
Sujith Pulikodan, Nihar Desai, Prasanta Kumar Ghosh
Thousands of languages are spoken worldwide, yet many remain under-resourced for Automatic Speech Recognition (ASR) due to the limited availability of high-quality transcribed spee…
eess.AS2026
Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction
Amartyaveer, Murali Kadambi, Chandra Mohan Sharma +2
In this study, we have presented a novel approach to predict the Short-Time Objective Intelligibility (STOI) metric using a bottleneck transformer architecture. Traditional methods…
eess.AS2025
Discovering phoneme-specific critical articulators through a data-driven approach
Jesuraj Bandekar, Sathvik Udupa, Prasanta Kumar Ghosh
We propose an approach for learning critical articulators for phonemes through a machine learning approach. We formulate the learning with three models trained end to end. First, w…