3 papers
cs.SD2026
Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization
Xiang Lin, Tian-Hao Zhang, Chunfeng Wang +3
Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is cha…
eess.AS2026
SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation
Haoyu Zhang, Yuta Oshima, Xingjian Du +4
Video-to-audio (V2A) generation must jointly satisfy audiovisual alignment, semantic consistency, temporal synchronization, and perceptual quality. While prior work has mainly focu…
eess.AS2023
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
Ziyue Jiang, Jinglin Liu, Yi Ren +10
Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping…