5 papers
LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction
Jinqing Zhang, Yanan Zhang, Wenrui Cai +2
Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong generalization to open-set objects…
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
Wenrui Cai, Zhenyi Lu, Yuzhe Li +4
With the advent of Transformer-based one-stream trackers that possess strong capability in inter-frame relation modeling, recent research has increasingly focused on how to introdu…
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
Jinqing Zhang, Zehua Fu, Zelin Xu +3
The comprehensive understanding capabilities of world models for driving scenarios have significantly improved the planning accuracy of end-to-end autonomous driving frameworks. Ho…
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
Yongchao Feng, Yajie Liu, Shuai Yang +13
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, th…
GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection
Jinqing Zhang, Yanan Zhang, Yunlong Qi +3
Bird's-Eye-View (BEV) representation has emerged as a mainstream paradigm for multi-view 3D object detection, demonstrating impressive perceptual capabilities. However, existing me…