activity
20222025
most citedBEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection

28 citations · 53 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Qiuchen Wang, Ruixue Ding, Zehui Chen +4

Understanding information from visually rich documents remains a significant challenge for traditional Retrieval-Augmented Generation (RAG) methods. Existing benchmarks predominant…

cs.CV2024

Point-DETR3D: Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection

Hongzhi Gao, Zheng Chen, Zehui Chen +4

Training high-accuracy 3D detectors necessitates massive labeled 3D annotations with 7 degree-of-freedom, which is laborious and time-consuming. Therefore, the form of point annota…

cs.CV2024

A Vanilla Multi-Task Framework for Dense Visual Prediction Solution to 1st VCL Challenge -- Multi-Task Robustness Track

Zehui Chen, Qiuchen Wang, Zhenyu Li +3

In this report, we present our solution to the multi-task robustness track of the 1st Visual Continual Learning (VCL) Challenge at ICCV 2023 Workshop. We propose a vanilla framewor…

cs.CV2024

Stream Query Denoising for Vectorized HD Map Construction

Shuo Wang, Fan Jia, Yingfei Liu +6

To enhance perception performance in complex and extensive scenarios within the realm of autonomous driving, there has been a noteworthy focus on temporal modeling, with a particul…

cs.CV2023

Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View

Shuo Wang, Xinhai Zhao, Hai-Ming Xu +5

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D o…

cs.CV2022★ 28 cited

BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection

Zehui Chen, Zhenyu Li, Shiquan Zhang +3

3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. Owing to its low cost and high efficiency, multi-view 3D object…