2 papers
cs.SD2026
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jun Zhan, Chen Yang, Yitian Gong +23
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…
cs.CL2025
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
Xingjian Zhao, Zhe Xu, Qinyuan Cheng +20
Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits exp…