3 papers
cs.RO2026
AgentVLN: Towards Agentic Vision-and-Language Navigation
Zihao Xin, Wentong Li, Yixuan Jiang +6
Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-La…
cs.CL2026
History-Guided Iterative Visual Reasoning with Self-Correction
Xinglong Yang, Zhilin Peng, Zhanzhan Liu +2
Self-consistency methods are the core technique for improving the reasoning reliability of multimodal large language models (MLLMs). By generating multiple reasoning results throug…
cs.CV2025
Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition
Pei Peng, MingKun Xie, Hang Hao +2
Object-context shortcuts remain a persistent challenge in vision-language models, undermining zero-shot reliability when test-time scenes differ from familiar training co-occurrenc…