Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
Sandro Papais, Lezhou Feng, Charles Cossette +1
Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across multiple views and large 3D…
cs.CV2025
DuoSpaceNet: Leveraging Both Bird's-Eye-View and Perspective View Representations for 3D Object Detection
Zhe Huang, Yizhe Zhao, Hao Xiao +2
Multi-view camera-only 3D object detection largely follows two primary paradigms: exploiting bird's-eye-view (BEV) representations or focusing on perspective-view (PV) features, ea…
cs.CV2024
MGTR: Multi-Granular Transformer for Motion Prediction with LiDAR
Yiqian Gan, Hao Xiao, Yizhe Zhao +4
Motion prediction has been an essential component of autonomous driving systems since it handles highly uncertain and complex scenarios involving moving agents of different types.…