95 citations · 128 across the 7 of their papers we have counts for
6 papers
Advancing Vision Transformers with Group-Mix Attention
Chongjian Ge, Xiaohan Ding, Zhan Tong +4
Vision Transformers (ViTs) have been shown to enhance visual recognition through modeling long-range dependencies with multi-head self-attention (MHSA), which is typically formulat…
Large Language Models as Automated Aligners for benchmarking Vision-Language Models
Yuanfeng Ji, Chongjian Ge, Weikai Kong +4
With the advancements in Large Language Models (LLMs), Vision-Language Models (VLMs) have reached a new level of sophistication, showing notable competence in executing intricate c…
Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
Youwei Liang, Chongjian Ge, Zhan Tong +3
Vision Transformers (ViTs) take all the image patches as tokens and construct multi-head self-attention (MHSA) among them. Complete leverage of these image tokens brings redundant…
Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning
Chongjian Ge, Youwei Liang, Yibing Song +3
Studies on self-supervised visual representation learning (SSL) improve encoder backbones to discriminate training samples without labels. While CNN encoders via SSL achieve compar…
Disentangled Cycle Consistency for Highly-realistic Virtual Try-On
Chongjian Ge, Yibing Song, Yuying Ge +3
Image virtual try-on replaces the clothes on a person image with a desired in-shop clothes image. It is challenging because the person and the in-shop clothes are unpaired. Existin…
Parser-Free Virtual Try-on via Distilling Appearance Flows
Yuying Ge, Yibing Song, Ruimao Zhang +3
Image virtual try-on aims to fit a garment image (target clothes) to a person image. Prior methods are heavily based on human parsing. However, slightly-wrong segmentation results…