20 citations · 49 across the 16 of their papers we have counts for
9 papers
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Lin Chen, Jinsong Li, Xiaoyi Dong +5
In the realm of large multi-modal models (LMMs), efficient modality alignment is crucial yet often constrained by the scarcity of high-quality image-text data. To address this bott…
Ultra-High Resolution Segmentation with Ultra-Rich Context: A Novel Benchmark
Deyi Ji, Feng Zhao, Hongtao Lu +2
With the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-g…
Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View
Shuo Wang, Xinhai Zhao, Hai-Ming Xu +5
Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D o…
CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition
Hanting Li, Hongjing Niu, Zhaoqing Zhu +1
Facial expression recognition (FER) is an essential task for understanding human behaviors. As one of the most informative behaviors of humans, facial expressions are often compoun…
Intensity-Aware Loss for Dynamic Facial Expression Recognition in the Wild
Hanting Li, Hongjing Niu, Zhaoqing Zhu +1
Compared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natu…
AutoAlignV2: Deformable Feature Aggregation for Dynamic Multi-Modal 3D Object Detection
Zehui Chen, Zhenyu Li, Shiquan Zhang +3
Point clouds and RGB images are two general perceptional sources in autonomous driving. The former can provide accurate localization of objects, and the latter is denser and richer…