activity
20242026
collaborators

26 papers

cs.CV2026

MOSS-VL Technical Report

Pengyu Wang, Chenkun Tan, Shaojun Zhou +29

We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across th…

cs.SD2026

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Jun Zhan, Chen Yang, Yitian Gong +23

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly ge…

cs.SD2026

MOSS Transcribe Diarize Technical Report

MOSI. AI, :, Donghua Yu +23

Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…

cs.RO2026

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

Li Ji, Siyin Wang, Pengfang Qian +5

Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance…

cs.RO2026

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

Junhao Shi, Zezheng Huai, Siyin Wang +7

Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, na…

cs.SD2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen +27

MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…