19 papers
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Senyu Fei, Xiaopeng Yu, Siyin Wang +3
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely o…
MOSS Transcribe Diarize Technical Report
MOSI. AI, :, Donghua Yu +23
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
Li Ji, Siyin Wang, Pengfang Qian +5
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance…
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
Junhao Shi, Siyin Wang, Xiaopeng Yu +3
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly t…
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
Yitian Gong, Kuangwei Chen, Zhaoye Fei +9
Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches…
MOVA: Towards Scalable and Synchronized Video-Audio Generation
OpenMOSS Team, Donghua Yu, Mingshu Chen +38
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…