collaborators

7 papers

cs.SD2026

MOSS Transcribe Diarize Technical Report

MOSI. AI, :, Donghua Yu +23

Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…

cs.SD2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen +27

MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…

cs.GR2026

MotionDuet: Dual-Conditioned 3D Human Motion Generation with Video-Regularized Text Learning

Yi-Yang Zhang, Tengjiao Sun, Pengcheng Fang +4

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work…

cs.SD2026

MOSS-TTSD: Text to Spoken Dialogue Generation

Yuqian Zhang, Donghua Yu, Zhengyuan Lin +15

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance t…

cs.SD2026

MOSS-TTS Technical Report

Yitian Gong, Botian Jiang, Yiwei Zhao +23

This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…

cs.CV2026

MOVA: Towards Scalable and Synchronized Video-Audio Generation

OpenMOSS Team, Donghua Yu, Mingshu Chen +38

Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…