activity
20222024
most citedAutoAlignV2: Deformable Feature Aggregation for Dynamic Multi-Modal 3D Object Detection

20 citations · 49 across the 16 of their papers we have counts for

collaborators

9 papers

cs.CV202312 cited

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Lin Chen, Jinsong Li, Xiaoyi Dong +5

In the realm of large multi-modal models (LMMs), efficient modality alignment is crucial yet often constrained by the scarcity of high-quality image-text data. To address this bott…

cs.CV2023

Ultra-High Resolution Segmentation with Ultra-Rich Context: A Novel Benchmark

Deyi Ji, Feng Zhao, Hongtao Lu +2

With the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-g…

cs.CV2023

Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View

Shuo Wang, Xinhai Zhao, Hai-Ming Xu +5

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D o…

cs.CV20234 cited

CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition

Hanting Li, Hongjing Niu, Zhaoqing Zhu +1

Facial expression recognition (FER) is an essential task for understanding human behaviors. As one of the most informative behaviors of humans, facial expressions are often compoun…

cs.CV20225 cited

Intensity-Aware Loss for Dynamic Facial Expression Recognition in the Wild

Hanting Li, Hongjing Niu, Zhaoqing Zhu +1

Compared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natu…

cs.CV202220 cited

AutoAlignV2: Deformable Feature Aggregation for Dynamic Multi-Modal 3D Object Detection

Zehui Chen, Zhenyu Li, Shiquan Zhang +3

Point clouds and RGB images are two general perceptional sources in autonomous driving. The former can provide accurate localization of objects, and the latter is denser and richer…