5 papers
STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs
Kaiyuan Zhang, Mohan Shi, Eray Eren +3
Neural audio codecs are widely used for audio compression and can be integrated into token-based language models. Traditional codecs preserve acoustic details well but lack semanti…
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
Mohan Shi, Natarajan Balaji Shankar, Kaiyuan Zhang +2
Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semanti…
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
Mohan Shi, Xiong Xiao, Ruchao Fan +2
Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end spe…
Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet
Anyu Ying, Natarajan Balaji Shankar, Chyi-Jiunn Lin +7
Despite advancements in ASR, child speech recognition remains challenging due to acoustic variability and limited annotated data. While fine-tuning adult ASR models on child speech…
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang +2
Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, lea…