activity
20182025
most citedMotion Transformer with Global Intention Localization and Local Movement Refinement

74 citations · 307 across the 30 of their papers we have counts for

collaborators
Showing 2023Show all

7 papers · 1 filter

cs.CV2023★ 1 cited

UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View Representation

Haiyang Wang, Hao Tang, Shaoshuai Shi +4

Jointly processing information from multiple sensors is crucial to achieving accurate and robust perception for reliable autonomous driving systems. However, current 3D perception…

cs.CV2023★ 4 cited

MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying

Shaoshuai Shi, Li Jiang, Dengxin Dai +1

Motion prediction is crucial for autonomous driving systems to understand complex driving scenarios and make informed decisions. However, this task is challenging due to the divers…

cs.CV2023

TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypotheses

Xuesong Chen, Shaoshuai Shi, Chao Zhang +5

3D multi-object tracking (MOT) is vital for many applications including autonomous driving vehicles and service robots. With the commonly used tracking-by-detection paradigm, 3D MO…

cs.CV2023★ 1 cited

Self-supervised Pre-training with Masked Shape Prediction for 3D Scene Understanding

Li Jiang, Zetong Yang, Shaoshuai Shi +3

Masked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this p…

cs.CV2023

Sparse Dense Fusion for 3D Object Detection

Yulu Gao, Chonghao Sima, Shaoshuai Shi +3

With the prevalence of multimodal learning, camera-LiDAR fusion has gained popularity in 3D object detection. Although multiple fusion approaches have been proposed, they can be cl…

cs.CV2023★ 11 cited

Virtual Sparse Convolution for Multimodal 3D Object Detection

Hai Wu, Chenglu Wen, Shaoshuai Shi +2

Recently, virtual/pseudo-point-based 3D object detection that seamlessly fuses RGB images and LiDAR data by depth completion has gained great attention. However, virtual points gen…