collaborators

6 papers

eess.AS2026

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

Zilai Wang, Natarajan Balaji Shankar, Mohan Shi +2

Speech foundation models often struggle in low-resource domains due to domain mismatch and data scarcity. We propose Gumbel-BEARD, a domain adaptation framework that automates Whis…

eess.AS2026

GC-LoRA: Gated Convolutional LoRA for Parameter-Efficient Acoustic Adaptation

Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang +2

Transformer-based Speech Foundation Models excel in most Automatic Speech Recognition tasks but often suffer performance degradation when applied to domains with mismatched acousti…

eess.AS2026

Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR

Mohan Shi, Kaiyuan Zhang, Zilai Wang +3

While Speech Large Language Models (Speech-LLMs) have achieved strong performance on adult Automatic Speech Recognition (ASR), their effectiveness on child speech remains under-exp…

eess.AS2026

STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs

Kaiyuan Zhang, Mohan Shi, Eray Eren +3

Neural audio codecs are widely used for audio compression and can be integrated into token-based language models. Traditional codecs preserve acoustic details well but lack semanti…

eess.AS2025

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs

Eray Eren, Qingju Liu, Hyeongwoo Kim +2

Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic feature…

eess.AS2025

G-IFT: A Gated Linear Unit adapter with Iterative Fine-Tuning for Low-Resource Children's Speaker Verification

Vishwas M. Shetty, Jiusi Zheng, Abeer Alwan

Speaker Verification (SV) systems trained on adults speech often underperform on children's SV due to the acoustic mismatch, and limited children speech data makes fine-tuning not…