collaborators

5 papers

eess.AS2026

STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs

Kaiyuan Zhang, Mohan Shi, Eray Eren +3

Neural audio codecs are widely used for audio compression and can be integrated into token-based language models. Traditional codecs preserve acoustic details well but lack semanti…

cs.CL2026

Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR

Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang +2

Self-supervised learning (SSL) models have achieved impressive results across many speech tasks, yet child automatic speech recognition (ASR) remains challenging due to limited dat…

eess.AS2025

Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR

Mohan Shi, Natarajan Balaji Shankar, Kaiyuan Zhang +2

Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semanti…

cs.CL2025

Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness

Lingnan Xu, Chong Feng, Kaiyuan Zhang +3

While large language models (LLMs) demonstrate impressive capabilities, their reliance on parametric knowledge often leads to factual inaccuracies. Retrieval-Augmented Generation (…

eess.AS2025

CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR

Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang +2

Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, lea…