From the 1 of 20 linked papers with an AI index.
11 papers · 1 filter
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
Jie Zhang, Xiaoyue Chen, Anzhe Chen +36
We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically ground…
Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation
Xunyi Zhao, Sihao Lin, Gengze Zhou +5
Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an under-specified natural-language description. Such ambiguity…
LightMover: Generative Light Movement with Color and Intensity Controls
Gengze Zhou, Tianyu Wang, Soo Ye Kim +7
We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes w…
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
Zicheng Duan, Jiatong Xia, Zeyu Zhang +7
Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implici…
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
Songze Li, Zun Wang, Gengze Zhou +8
Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instruction…
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
Xunyi Zhao, Gengze Zhou, Qi Wu
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, whic…