collaborators

7 papers

cs.SD2026

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Jun Zhan, Chen Yang, Yitian Gong +23

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…

cs.SD2026

MOSS Transcribe Diarize Technical Report

MOSI. AI, :, Donghua Yu +23

Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…

cs.SD2026

MOSS-TTS Technical Report

Yitian Gong, Botian Jiang, Yiwei Zhao +23

This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…

cs.SD2026

MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models

Yitian Gong, Kuangwei Chen, Zhaoye Fei +9

Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches…

cs.CL2025

MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance

Xingjian Zhao, Zhe Xu, Qinyuan Cheng +20

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits exp…

eess.AS2025

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

Ruifan Deng, Yitian Gong, Qinghui Gao +5

With the rise of multimodal large language models (LLMs), audio codec plays an increasingly vital role in encoding audio into discrete tokens, enabling integration of audio into te…