3 papers
cs.CV2026
MOSS-VL Technical Report
Pengyu Wang, Chenkun Tan, Shaojun Zhou +29
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across th…
cs.CV2026
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention
Pengyu Wang, Chenkun Tan, Shaojun Zhou +18
Video understanding is shifting from the offline paradigm -- taking a fully recorded video as input and producing a single answer after it ends -- toward real-time interaction, in…
cs.AI2026
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation
Zixuan Jiang, Yanqiao Zhu, Peng Wang +8
Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most cur…