26 papers
MOSS-VL Technical Report
Pengyu Wang, Chenkun Tan, Shaojun Zhou +29
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across th…
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jun Zhan, Chen Yang, Yitian Gong +23
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…
MOSS Transcribe Diarize Technical Report
MOSI. AI, :, Donghua Yu +23
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
Li Ji, Siyin Wang, Pengfang Qian +5
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance…
Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy
Junhao Shi, Zezheng Huai, Siyin Wang +7
Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, na…
MOSS-Audio Technical Report
Chen Yang, Chufan Yu, Hanfu Chen +27
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…