6 papers · 1 filter
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
Mohan Shi, Natarajan Balaji Shankar, Kaiyuan Zhang +2
Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semanti…
An Age-Agnostic System for Robust Speaker Verification
Jiusi Zheng, Vishwas Shetty, Natarajan Balaji Shankar +1
In speaker verification (SV), the acoustic mismatch between children's and adults' speech leads to suboptimal performance when adult-trained SV systems are applied to children's sp…
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang +2
Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, lea…
Enhancing Age-Related Robustness in Children Speaker Verification
Vishwas M. Shetty, Jiusi Zheng, Steven M. Lulich +1
One of the main challenges in children's speaker verification (C-SV) is the significant change in children's voices as they grow. In this paper, we propose two approaches to improv…
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
Natarajan Balaji Shankar, Ruchao Fan, Abeer Alwan
Recently, speech foundation models have gained popularity due to their superiority in finetuning downstream ASR tasks. However, models finetuned on certain domains, such as LibriSp…
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
Ruchao Fan, Natarajan Balaji Shankar, Abeer Alwan
Speech foundation models (SFMs) have achieved state-of-the-art results for various speech tasks in supervised (e.g. Whisper) or self-supervised systems (e.g. WavLM). However, the p…