activity
20172023
most citedShuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer

125 citations · 237 across the 13 of their papers we have counts for

collaborators

21 papers

cs.CV202317 cited

BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Yang Zhao, Zhijie Lin, Daquan Zhou +3

LLMs have demonstrated remarkable abilities at interacting with humans through language, especially with the usage of instruction-following data. Recent advancements in LLMs, such…

cs.CV2023

Disentangled Pre-training for Image Matting

Yanda Li, Zilong Huang, Gang Yu +3

Image matting requires high-quality pixel-level human annotations to support the training of a deep model in recent literature. Whereas such annotation is costly and hard to scale,…

cs.CV2022

Efficient Single-Image Depth Estimation on Mobile Devices, Mobile AI & AIM 2022 Challenge: Report

Andrey Ignatov, Grigory Malivenko, Radu Timofte +36

Various depth estimation models are now widely used on many mobile and IoT devices for image segmentation, bokeh effect rendering, object tracking and many other mobile tasks. Thus…

cs.CV20226 cited

Coordinates Are NOT Lonely -- Codebook Prior Helps Implicit Neural 3D Representations

Fukun Yin, Wen Liu, Zilong Huang +3

Implicit neural 3D representation has achieved impressive results in surface or scene reconstruction and novel view synthesis, which typically uses the coordinate-based multi-layer…

cs.CV202216 cited

TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation

Wenqiang Zhang, Zilong Huang, Guozhong Luo +5

Although vision transformers (ViTs) have achieved great success in computer vision, the heavy computational cost hampers their applications to dense prediction tasks such as semant…

cs.CV20211 cited

Shuffle Transformer with Feature Alignment for Video Face Parsing

Rui Zhang, Yang Han, Zilong Huang +4

This is a short technical report introducing the solution of the Team TCParser for Short-video Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR…