activity
20172023
most citedTransReID: Transformer-based Object Re-Identification

95 citations · 211 across the 22 of their papers we have counts for

collaborators
Showing cs.CVShow all

27 papers · 1 filter

cs.CV20232 cited

Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition

Yujun Ma, Benjia Zhou, Ruili Wang +1

RGB-D action and gesture recognition remain an interesting topic in human-centered scene understanding, primarily due to the multiple granularities and large variation in human mot…

cs.CV2023

SCT: A Simple Baseline for Parameter-Efficient Fine-Tuning via Salient Channels

Henry Hengyuan Zhao, Pichao Wang, Yuyang Zhao +3

Pre-trained vision transformers have strong representation benefits to various downstream tasks. Recently, many parameter-efficient fine-tuning (PEFT) methods have been proposed, a…

cs.CV20231 cited

Revisiting Vision Transformer from the View of Path Ensemble

Shuning Chang, Pichao Wang, Hao Luo +2

Vision Transformers (ViTs) are normally regarded as a stack of transformer layers. In this work, we propose a novel view of ViTs showing that they can be seen as ensemble networks…

cs.CV2023

Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

Sarah Ibrahimi, Xiaohang Sun, Pichao Wang +3

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods…

cs.CV2023

DOAD: Decoupled One Stage Action Detection Network

Shuning Chang, Pichao Wang, Fan Wang +2

Localizing people and recognizing their actions from videos is a challenging task towards high-level video understanding. Existing methods are mostly two-stage based, with one stag…

cs.CV20236 cited

PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation

Qitao Zhao, Ce Zheng, Mengyuan Liu +2

Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relation…