4 papers
MOSS-Audio Technical Report
Chen Yang, Chufan Yu, Hanfu Chen +27
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
Sihang Zhao, Kangrui Yu, Youliang Yuan +2
Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where stu…
MOSS-TTS Technical Report
Yitian Gong, Botian Jiang, Yiwei Zhao +23
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretrainin…
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
Yuanhe Guo, Linxi Xie, Zhuoran Chen +5
We introduce ImageGem, a dataset for studying generative models that understand fine-grained individual preferences. We posit that a key challenge hindering the development of such…