Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
MOSS-Audio Technical Report
Chen Yang, Chufan Yu, Hanfu Chen +27
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…
cs.SD2026
MOSS-TTS Technical Report
Yitian Gong, Botian Jiang, Yiwei Zhao +23
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…
cs.SD2026
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
Yitian Gong, Kuangwei Chen, Zhaoye Fei +9
Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches…