collaborators

10 papers

cs.SD2026

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

Yuqian Zhang, Yao Shi, Kexin Huang +6

Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent s…

cs.SD2026

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Jun Zhan, Chen Yang, Yitian Gong +23

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…

cs.SD2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen +27

MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…

cs.SD2026

MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions

Kexin Huang, Liwei Fan, Botian Jiang +11

Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, person…

cs.SD2026

MOSS-TTS Technical Report

Yitian Gong, Botian Jiang, Yiwei Zhao +23

This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…

cs.SD2026

MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models

Yitian Gong, Kuangwei Chen, Zhaoye Fei +9

Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches…