Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MUSE: A Unified Agentic Harness for MLLMs
Jianglin Lu, Hailing Wang, Xu Ma +4
Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze from a screenshot or selecting t…
cs.CV2026
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
Chufan Shi, Cheng Yang, Yaokang Wu +4
Vision-Language Models (VLMs) often produce self-reflective statements like "let me check the figure again" during reasoning. Do such statements trigger genuine visual re-examinati…