12 citations · 20 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 12 cited
LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
Qiang Chen, Xiangbo Su, Xinyu Zhang +12
In this paper, we present a light-weight detection transformer, LW-DETR, which outperforms YOLOs for real-time object detection. The architecture is a simple stack of a ViT encoder…
cs.CV2023★ 5 cited
HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
Junkun Yuan, Xinyu Zhang, Hao Zhou +12
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting…
cs.CV2023★ 3 cited
Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identification
Zhiyin Shao, Xinyu Zhang, Changxing Ding +2
The pre-training task is indispensable for the text-to-image person re-identification (T2I-ReID) task. However, there are two underlying inconsistencies between these two tasks tha…