4 papers
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
Ruihao Jing, Cheng Gong, Yu Jiang +5
Rare words remain a critical bottleneck for speech-to-text systems. While direct fine-tuning improves recognition of target words, it often incurs high cost, catastrophic forgettin…
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Xueqing Li, Hao Ma, Zehan Li +8
Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…
: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
Boyu Zhu, Cheng Gong, Muyang Wu +5
Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these mode…
AudioSpa: Spatializing Sound Events with Text
Linfeng Feng, Lei Zhao, Boyu Zhu +2
Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text…