activity
20212025
most citedNot All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

95 citations · 128 across the 7 of their papers we have counts for

collaborators

6 papers

cs.CV202312 cited

Advancing Vision Transformers with Group-Mix Attention

Chongjian Ge, Xiaohan Ding, Zhan Tong +4

Vision Transformers (ViTs) have been shown to enhance visual recognition through modeling long-range dependencies with multi-head self-attention (MHSA), which is typically formulat…

cs.CV2023

Large Language Models as Automated Aligners for benchmarking Vision-Language Models

Yuanfeng Ji, Chongjian Ge, Weikai Kong +4

With the advancements in Large Language Models (LLMs), Vision-Language Models (VLMs) have reached a new level of sophistication, showing notable competence in executing intricate c…

cs.CV202295 cited

Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Youwei Liang, Chongjian Ge, Zhan Tong +3

Vision Transformers (ViTs) take all the image patches as tokens and construct multi-head self-attention (MHSA) among them. Complete leverage of these image tokens brings redundant…

cs.CV20216 cited

Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning

Chongjian Ge, Youwei Liang, Yibing Song +3

Studies on self-supervised visual representation learning (SSL) improve encoder backbones to discriminate training samples without labels. While CNN encoders via SSL achieve compar…

cs.CV20214 cited

Disentangled Cycle Consistency for Highly-realistic Virtual Try-On

Chongjian Ge, Yibing Song, Yuying Ge +3

Image virtual try-on replaces the clothes on a person image with a desired in-shop clothes image. It is challenging because the person and the in-shop clothes are unpaired. Existin…

cs.CV202111 cited

Parser-Free Virtual Try-on via Distilling Appearance Flows

Yuying Ge, Yibing Song, Ruimao Zhang +3

Image virtual try-on aims to fit a garment image (target clothes) to a person image. Prior methods are heavily based on human parsing. However, slightly-wrong segmentation results…