4 citations · 12 across the 9 of their papers we have counts for
9 papers
BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection
Wenjie Wang, Yehao Lu, Guangcong Zheng +6
Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it encompasses inherent advantages in reducing blind spots and expandin…
Training-Free Unsupervised Prompt for Vision-Language Models
Sifan Long, Linbin Wang, Zhen Zhao +4
Prompt learning has become the most effective paradigm for adapting large pre-trained vision-language models (VLMs) to downstream tasks. Recently, unsupervised prompt tuning method…
PVLR: Prompt-driven Visual-Linguistic Representation Learning for Multi-Label Image Recognition
Hao Tan, Zichang Tan, Jun Li +2
Multi-label image recognition is a fundamental task in computer vision. Recently, vision-language models have made notable advancements in this area. However, previous methods ofte…
ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification
Guiwei Zhang, Yongfei Zhang, Zichang Tan
Visible-infrared person re-identification is challenging due to the large modality gap. To bridge the gap, most studies heavily rely on the correlation of visible-infrared holistic…
Unified Frequency-Assisted Transformer Framework for Detecting and Grounding Multi-Modal Manipulation
Huan Liu, Zichang Tan, Qiang Chen +3
Detecting and grounding multi-modal media manipulation (DGM^4) has become increasingly crucial due to the widespread dissemination of face forgery and text misinformation. In this…
Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation
Huan Liu, Qiang Chen, Zichang Tan +9
In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.…