4 papers
GSRM: Generative Speech Reward Model for Speech RLHF
Maohao Shen, Tejas Jayashankar, Osama Hanna +10
Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…
Scaling Speech Tokenizers with Diffusion Autoencoders
Yuancheng Wang, Zhenyu Tang, Yun Wang +9
Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understandi…
VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
Yancheng Wang, Osama Hanna, Ruiming Xie +11
Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features suc…
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
Shivam Mehta, Yingru Liu, Zhenyu Tang +6
Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker…