4 papers
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jun Zhan, Chen Yang, Yitian Gong +23
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…
Switchcodec: Adaptive residual-expert sparse quantization for high-fidelity neural audio coding
Xiangbo Wang, Wenbin Jiang, Jin Wang +3
Recent neural audio compression models often rely on residual vector quantization for high-fidelity coding, but using a fixed number of per-frame codebooks is suboptimal for the wi…
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
Jin Wang, Wenbin Jiang, Xiangbo Wang +2
Neural audio compression has emerged as a promising technology for efficiently representing speech, music, and general audio. However, existing methods suffer from significant perf…
MOSS-TTS Technical Report
Yitian Gong, Botian Jiang, Yiwei Zhao +23
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…