most citedFM-ViT: Flexible Modal Vision Transformers for Face Anti-Spoofing

4 citations · 12 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV20241 cited

BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection

Wenjie Wang, Yehao Lu, Guangcong Zheng +6

Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it encompasses inherent advantages in reducing blind spots and expandin…

cs.CV2024

Training-Free Unsupervised Prompt for Vision-Language Models

Sifan Long, Linbin Wang, Zhen Zhao +4

Prompt learning has become the most effective paradigm for adapting large pre-trained vision-language models (VLMs) to downstream tasks. Recently, unsupervised prompt tuning method…

cs.CV20241 cited

PVLR: Prompt-driven Visual-Linguistic Representation Learning for Multi-Label Image Recognition

Hao Tan, Zichang Tan, Jun Li +2

Multi-label image recognition is a fundamental task in computer vision. Recently, vision-language models have made notable advancements in this area. However, previous methods ofte…

cs.CV2023

ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification

Guiwei Zhang, Yongfei Zhang, Zichang Tan

Visible-infrared person re-identification is challenging due to the large modality gap. To bridge the gap, most studies heavily rely on the correlation of visible-infrared holistic…

cs.CV20232 cited

Unified Frequency-Assisted Transformer Framework for Detecting and Grounding Multi-Modal Manipulation

Huan Liu, Zichang Tan, Qiang Chen +3

Detecting and grounding multi-modal media manipulation (DGM^4) has become increasingly crucial due to the widespread dissemination of face forgery and text misinformation. In this…

cs.CV20233 cited

Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation

Huan Liu, Qiang Chen, Zichang Tan +9

In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.…