3 papers
cs.CV2026
SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
Sandro Papais, Lezhou Feng, Charles Cossette +1
Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across multiple views and large 3D…
cs.LG2026
Learning Dynamic Representations via An Optimally-Weighted Maximum Mean Discrepancy Optimization Framework for Continual Learning
KaiHui Huang, RunQing Wu, JinHui Sheng +4
Continual learning has emerged as a pivotal area of research, primarily due to its advantageous characteristic that allows models to persistently acquire and retain information. Ho…
cs.CV2025
DuoSpaceNet: Leveraging Both Bird's-Eye-View and Perspective View Representations for 3D Object Detection
Zhe Huang, Yizhe Zhao, Hao Xiao +2
Multi-view camera-only 3D object detection largely follows two primary paradigms: exploiting bird's-eye-view (BEV) representations or focusing on perspective-view (PV) features, ea…