85 citations · 133 across the 6 of their papers we have counts for
5 papers · 1 filter
TCFormer: Visual Recognition via Token Clustering Transformer
Wang Zeng, Sheng Jin, Lumin Xu +5
Transformers are widely used in computer vision areas and have achieved remarkable success. Most state-of-the-art approaches split images into regular grids and represent each grid…
Unlock the Power: Competitive Distillation for Multi-Modal Large Language Models
Xinwei Li, Li Lin, Shuai Wang +1
Recently, multi-modal content generation has attracted lots of attention from researchers by investigating the utilization of visual instruction tuning based on large language mode…
Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
Junhong Gou, Siyu Sun, Jianfu Zhang +3
Virtual try-on is a critical image synthesis task that aims to transfer clothes from one image to another while preserving the details of both humans and clothes. While many existi…
ZoomNAS: Searching for Whole-body Human Pose Estimation in the Wild
Lumin Xu, Sheng Jin, Wentao Liu +4
This paper investigates the task of 2D whole-body human pose estimation, which aims to localize dense landmarks on the entire human body including body, feet, face, and hands. We p…
Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary Detection
Jiaqi Tang, Zhaoyang Liu, Chen Qian +2
Generic event boundary detection is an important yet challenging task in video understanding, which aims at detecting the moments where humans naturally perceive event boundaries.…