10 papers
Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains
Zilai Wang, Natarajan Balaji Shankar, Mohan Shi +2
Speech foundation models often struggle in low-resource domains due to domain mismatch and data scarcity. We propose Gumbel-BEARD, a domain adaptation framework that automates Whis…
GC-LoRA: Gated Convolutional LoRA for Parameter-Efficient Acoustic Adaptation
Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang +2
Transformer-based Speech Foundation Models excel in most Automatic Speech Recognition tasks but often suffer performance degradation when applied to domains with mismatched acousti…
Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR
Mohan Shi, Kaiyuan Zhang, Zilai Wang +3
While Speech Large Language Models (Speech-LLMs) have achieved strong performance on adult Automatic Speech Recognition (ASR), their effectiveness on child speech remains under-exp…
STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs
Kaiyuan Zhang, Mohan Shi, Eray Eren +3
Neural audio codecs are widely used for audio compression and can be integrated into token-based language models. Traditional codecs preserve acoustic details well but lack semanti…
Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang +2
Self-supervised learning (SSL) models have achieved impressive results across many speech tasks, yet child automatic speech recognition (ASR) remains challenging due to limited dat…
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
Mohan Shi, Natarajan Balaji Shankar, Kaiyuan Zhang +2
Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semanti…