collaborators

6 papers

cs.RO2026

PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets

Jie Huang, Xiaohe Li, Jiahao Li +6

Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model that emits a serialized asset f…

cs.AI2026

Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities

Xiaohe Li, Yiru Wang, Junhao Fan +6

Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a…

cs.RO2026

Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

Jiazhao Zhang, Gengze Zhou, Hale Yin +32

Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search…

cs.CV2026

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Zekai Zhang, Jiahao Li, Jie Zhang +18

While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent on up-to-date knowl…

cs.CV2026

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

Jie Zhang, Xiaoyue Chen, Anzhe Chen +36

We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically ground…

cs.CV2025

EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents

Zhili Cheng, Yuge Tu, Ran Li +9

Multimodal Large Language Models (MLLMs) have shown significant advancements, providing a promising future for embodied agents. Existing benchmarks for evaluating MLLMs primarily u…