2 papers
cs.CV2026
Beyond Appearance: Can Multimodal Large Language Models Exploit Vertical Structure for Remote Sensing Natural Scene Understanding?
Jing Huang, Duanchu Wang, Junjie Yang +5
Multimodal large language models (MLLMs) have advanced rapidly in remote-sensing analysis, yet existing evaluations remain predominantly 2D-centric. Because spectrally confused reg…
cs.RO2026
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
Yayun He, Zuheng Kang, Botao Zhao +3
Vision-language models (VLMs) have significantly improved the generalization capabilities of robotic manipulation. However, VLM-based systems often suffer from a lack of robustness…