2 papers
cs.RO2026
S-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
Zhipeng Xie, Zongyi Han, Xiangyi Wei +3
Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulat…
cs.CV2026
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
Zhuang Yu, Lei Shen, Jing Zhao +1
Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, an…