2 papers
cs.AI2026
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
Wencheng Ye, Yi Bin, Yujuan Ding +7
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening eviden…
cs.RO2026
Language-Grounded Decoupled Action Representation for Robotic Manipulation
Wuding Weng, Tongshu Wu, Liucheng Chen +5
The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods hav…