76 citations · 147 across the 9 of their papers we have counts for
12 papers · 1 filter
Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation
Huan Liu, Qiang Chen, Zichang Tan +9
In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.…
Learning Structure-Guided Diffusion Model for 2D Human Pose Estimation
Zhongwei Qiu, Qiansheng Yang, Jian Wang +7
One of the mainstream schemes for 2D human pose estimation (HPE) is learning keypoints heatmaps by a neural network. Existing methods typically improve the quality of heatmaps by c…
Exploring Effective Factors for Improving Visual In-Context Learning
Yanpeng Sun, Qiang Chen, Xiaofan Li +3
The In-Context Learning (ICL) is to understand a new task via a few demonstrations (aka. prompt) and predict new inputs without tuning the models. While it has been widely studied…
CAE v2: Context Autoencoder with CLIP Target
Xinyu Zhang, Jiahui Chen, Junkun Yuan +10
Masked image modeling (MIM) learns visual representation by masking and reconstructing image patches. Applying the reconstruction supervision on the CLIP representation has been pr…
Group DETR v2: Strong Object Detector with Encoder-Decoder Pretraining
Qiang Chen, Jian Wang, Chuchu Han +12
We present a strong object detector with encoder-decoder pretraining and finetuning. Our method, called Group DETR v2, is built upon a vision transformer encoder ViT-Huge~\cite{dos…
U-HRNet: Delving into Improving Semantic Representation of High Resolution Network for Dense Prediction
Jian Wang, Xiang Long, Guowei Chen +3
High resolution and advanced semantic representation are both vital for dense prediction. Empirically, low-resolution feature maps often achieve stronger semantic representation, a…