8 papers
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
Fengyi Wu, Yifei Dong, Yilong Dai +7
We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentri…
FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI
Yuhang Peng, Yizhou Pan, Xinning He +6
As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex,…
RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
Xiyan Liu, Han Wang, Yuhu Wang +4
Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous…
Bench2FreeAD: A Benchmark for Vision-based End-to-end Navigation in Unstructured Robotic Environments
Yuhang Peng, Sidong Wang, Jihaoyu Yang +3
Most current end-to-end (E2E) autonomous driving algorithms are built on standard vehicles in structured transportation scenarios, lacking exploration of robot navigation for unstr…
Embodied Cognition Augmented End2End Autonomous Driving
Ling Niu, Xiaoji Zheng, Han Wang +4
In recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networ…