activity
20212023
most citedShuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer

125 citations · 179 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV202325 cited

ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Yucheng Han, Chi Zhang, Xin Chen +5

Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for spe…

cs.CV2022

Efficient Single-Image Depth Estimation on Mobile Devices, Mobile AI & AIM 2022 Challenge: Report

Andrey Ignatov, Grigory Malivenko, Radu Timofte +36

Various depth estimation models are now widely used on many mobile and IoT devices for image segmentation, bokeh effect rendering, object tracking and many other mobile tasks. Thus…

cs.CV20225 cited

Learning Variational Motion Prior for Video-based Motion Capture

Xin Chen, Zhuo Su, Lingbo Yang +4

Motion capture from a monocular video is fundamental and crucial for us humans to naturally experience and interact with each other in Virtual Reality (VR) and Augmented Reality (A…

cs.CV20228 cited

Hierarchical Normalization for Robust Monocular Depth Estimation

Chi Zhang, Wei Yin, Zhibin Wang +3

In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-th…

cs.CV20212 cited

Fine-grained Identity Preserving Landmark Synthesis for Face Reenactment

Haichao Zhang, Youcheng Ben, Weixi Zhang +3

Recent face reenactment works are limited by the coarse reference landmarks, leading to unsatisfactory identity preserving performance due to the distribution gap between the manip…

cs.CV20211 cited

Shuffle Transformer with Feature Alignment for Video Face Parsing

Rui Zhang, Yang Han, Zilong Huang +4

This is a short technical report introducing the solution of the Team TCParser for Short-video Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR…