collaborators

7 papers

eess.AS2026

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

Paban Sapkota, Hemant Kumar Kathania, Sudarsana Reddy Kadiri +1

Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria. However, accurately recognizing dysarthric speech poses signifi…

eess.AS2026

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo +2

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has d…

cs.SD2026

voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models

Aju Ani Justus, Ruchit Agrawal, Sudarsana Reddy Kadiri +1

We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervis…

eess.AS2025

Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?

Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri +1

Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. W…

eess.AS2025

Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech

Abhijit Sinha, Harishankar Kumar, Mohit Joshi +3

Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SS…

cs.LG2025

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

Sean Foley, Hong Nguyen, Jihwan Lee +4

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studie…