6 papers
Rethink Before You Execute: Adaptive Execution for World Action Models
Feng Ye, Yiming Zhao, Yong Yu +5
World Action Models (WAMs) jointly predict future actions and the evolution of the environment. At each inference, a WAM generates a chunk of actions and the robot executes a fixed…
RoamFlow: Reinforcement-Aligned One-Step Action MeanFlow Policy for Image-Goal Navigation
Zixuan Zhang, Yuqi Chen, Junjie Gao +4
Image-goal navigation is a key challenge in embodied robotics, where an agent must reach a target specified solely by a goal image. While existing reinforcement learning approaches…
Flying to Image-Specified Objects: 3D Quadrotor Navigation via Cross-Graph Memory and Viewpoint Planning
Junjie Gao, Yuqi Chen, Yongzhou Pan +3
Instance-Specific Image-Goal Navigation (InstanceImageNav) requires a robot to navigate toward the exact object instance depicted in a query image. Extending this task to quadrotor…
SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
Jiayuan Du, Yiming Zhao, Zhenglong Guo +5
This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs)…
QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
Sicheng Zuo, Wenzhao Zheng, Xiaoyong Han +3
3D occupancy prediction is crucial for robust autonomous driving systems as it enables comprehensive perception of environmental structures and semantics. Most existing methods emp…
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
Wenzhao Zheng, Junjie Wu, Yao Zheng +8
Vision-based autonomous driving shows great potential due to its satisfactory performance and low costs. Most existing methods adopt dense representations (e.g., bird's eye view) o…