7 papers
Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation
Paban Sapkota, Hemant Kumar Kathania, Sudarsana Reddy Kadiri +1
Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria. However, accurately recognizing dysarthric speech poses signifi…
Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models
Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo +2
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has d…
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
Aju Ani Justus, Ruchit Agrawal, Sudarsana Reddy Kadiri +1
We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervis…
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri +1
Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. W…
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
Abhijit Sinha, Harishankar Kumar, Mohit Joshi +3
Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SS…
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
Sean Foley, Hong Nguyen, Jihwan Lee +4
Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studie…