2 papers
cs.CV2026
VertiCue-Bench: Diagnosing Whether MLLMs Use Height Cues to Resolve 2D Ambiguity in Remote Sensing Natural Scenes
Jing Huang, Duanchu Wang, Junjie Yang +5
Multimodal Large Language Models (MLLMs) have recently shown promising progress in geospatial reasoning. However, existing remote sensing benchmarks remain largely 2D-centric, eval…
cs.RO2026
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
Yayun He, Zuheng Kang, Botao Zhao +3
Vision-language models (VLMs) have significantly improved the generalization capabilities of robotic manipulation. However, VLM-based systems often suffer from a lack of robustness…