works on

From the 1 of 20 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

Jie Zhang, Xiaoyue Chen, Anzhe Chen +36

We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically ground…

cs.CV2026

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation

Xunyi Zhao, Sihao Lin, Gengze Zhou +5

Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an under-specified natural-language description. Such ambiguity…

cs.CV2026

LightMover: Generative Light Movement with Color and Intensity Controls

Gengze Zhou, Tianyu Wang, Soo Ye Kim +7

We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes w…

cs.CV2026

LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models

Zicheng Duan, Jiatong Xia, Zeyu Zhang +7

Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implici…

cs.CV2026

Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale

Songze Li, Zun Wang, Gengze Zhou +8

Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instruction…

cs.CV2026

VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents

Xunyi Zhao, Gengze Zhou, Qi Wu

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, whic…