3 papers
cs.CV2026
Overcoming Statistical Bias in Action-Controllable World Models
Yuhong Shi, Zhenhao Chu, Jie Wei +3
Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet future frames are often highly predictable from visual inertia and recur…
cs.CV2026
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Kangyi Wu, Pengna Li, Kailin Lyu +5
Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Large Language Models(Video-LLM…
cs.CV2026
Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation
Zhen Liu, Yuhan Liu, Jinjun Wang +3
Vision-and-Language Navigation requires agents to follow natural-language instructions in visually changing environments. A central challenge is the dynamic entanglement between la…