7 papers
MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight
Zehua Fan, Junjie He, Wenxuan Song +14
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…
MambaBEV: An EV-based 3D detection model with Mamba2
Zihan You, Ni Wang, Hao Wang +2
Accurate 3D object detection in autonomous driving relies on Bird's Eye View (BEV) perception and effective temporal fusion. However, existing fusion strategies based on convolutio…
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
Zihan You, Hongwei Liu, Chenxu Dang +4
Recent advances in Vision-Language-Action (VLA) models have shown promising capabilities in autonomous driving by leveraging the understanding and reasoning strengths of Large Lang…
SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning
Chenxu Dang, Jie Wang, Guang Li +6
In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in indiv…
Hybrid guided variational autoencoder for visual place recognition
Ni Wang, Zihan You, Emre Neftci +1
Autonomous agents such as cars, robots and drones need to precisely localize themselves in diverse environments, including in GPS-denied indoor environments. One approach for preci…
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
Zhifei Chen, Tianshuo Xu, Leyi Wu +6
Video generation has recently made striking visual progress, but maintaining coherent object motion and interactions remains difficult. We trace two practical bottlenecks: (i) huma…