7 papers
Rethinking 3D Shape Generation: Diffusion over Superquadrics
Zhiyang Liu, Wanze Li, Yuwei Wu +4
Diffusion models have advanced 3D shape generation, yet most methods still denoise in high-cardinality spaces (e.g., voxel/SDF grids, meshes, or point clouds), which is computation…
NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving
Jiahui Li, Jiawei Sun, Zixiang Ren +7
Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compact scene tokens for downstre…
IMPACT: Behavioral Intention-aware Multimodal Trajectory Prediction with Adaptive Context Trimming
Jiawei Sun, Xibin Yue, Jiahui Li +6
While most prior research has focused on improving the precision of multimodal trajectory predictions, the explicit modeling of multimodal behavioral intentions (e.g., yielding, ov…
PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving
Chengran Yuan, Zijian Lu, Zhanqi Zhang +8
End-to-end motion planning is promising for simplifying complex autonomous driving pipelines. However, challenges such as scene understanding and effective prediction for decision-…
RMP-YOLO: A Robust Motion Predictor for Partially Observable Scenarios even if You Only Look Once
Jiawei Sun, Jiahui Li, Tingchen Liu +6
We introduce RMP-YOLO, a unified framework designed to provide robust motion predictions even with incomplete input data. Our key insight stems from the observation that complete a…
Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations
Yibo Cui, Liang Xie, Yu Zhao +2
Vision-Language Navigation (VLN) enables intelligent agents to navigate environments by integrating visual perception and natural language instructions, yet faces significant chall…