5 papers
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jun Zhan, Chen Yang, Yitian Gong +23
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…
MOSS-TTSD: Text to Spoken Dialogue Generation
Yuqian Zhang, Donghua Yu, Zhengyuan Lin +15
Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance t…
MOSS-TTS Technical Report
Yitian Gong, Botian Jiang, Yiwei Zhao +23
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…
RGAlign-Rec: Ranking-Guided Alignment for Latent Query Reasoning in Recommendation Systems
Junhua Liu, Yang Jihao, Cheng Chang +3
Proactive intent prediction is a critical capability in modern e-commerce chatbots, enabling "zero-query" recommendations by anticipating user needs from behavioral and contextual…
MOVA: Towards Scalable and Synchronized Video-Audio Generation
OpenMOSS Team, Donghua Yu, Mingshu Chen +38
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…