96 citations · 145 across the 8 of their papers we have counts for
9 papers
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Bohao Zhao, Chengrui Wei, Guangfeng Jiang +17
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive percept…
X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
Baolu Li, Jingyu Qian, Rui Guo +17
Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive…
Imitation with Spatial-Temporal Heatmap: 2nd Place Solution for NuPlan Challenge
Yihan Hu, Kun Li, Pingyuan Liang +7
This paper presents our 2nd place solution for the NuPlan Challenge 2023. Autonomous driving in real-world scenarios is highly complex and uncertain. Achieving safe planning in the…
AFDetV2: Rethinking the Necessity of the Second Stage for Object Detection from Point Clouds
Yihan Hu, Zhuangzhuang Ding, Runzhou Ge +4
There have been two streams in the 3D detection from point clouds: single-stage methods and two-stage methods. While the former is more computationally efficient, the latter usuall…
Real-Time Anchor-Free Single-Stage 3D Detection with IoU-Awareness
Runzhou Ge, Zhuangzhuang Ding, Yihan Hu +4
In this report, we introduce our winning solution to the Real-time 3D Detection and also the "Most Efficient Model" in the Waymo Open Dataset Challenges at CVPR 2021. Extended from…
AFDet: Anchor Free One Stage 3D Object Detection
Runzhou Ge, Zhuangzhuang Ding, Yihan Hu +4
High-efficiency point cloud 3D object detection operated on embedded systems is important for many robotics applications including autonomous driving. Most previous works try to so…