2 papers
cs.SD2026
Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization
Xiang Lin, Tian-Hao Zhang, Chunfeng Wang +3
Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is cha…
eess.AS2026
SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation
Haoyu Zhang, Yuta Oshima, Xingjian Du +4
Video-to-audio (V2A) generation must jointly satisfy audiovisual alignment, semantic consistency, temporal synchronization, and perceptual quality. While prior work has mainly focu…