activity
20202022
most citedReal-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data

85 citations · 174 across the 23 of their papers we have counts for

collaborators

27 papers

cs.CV2022

Darwinian Model Upgrades: Model Evolving with Selective Compatibility

Binjie Zhang, Shupeng Su, Yixiao Ge +5

The traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensiv…

cs.CV20225 cited

MonoNeuralFusion: Online Monocular Neural 3D Reconstruction with Geometric Priors

Zi-Xin Zou, Shi-Sheng Huang, Yan-Pei Cao +3

High-fidelity 3D scene reconstruction from monocular videos continues to be challenging, especially for complete and fine-grained geometry reconstruction. The previous 3D reconstru…

cs.CV20224 cited

Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection

Yuxin Fang, Shusheng Yang, Shijie Wang +3

We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…

cs.CV2022

Accelerating the Training of Video Super-Resolution Models

Lijian Lin, Xintao Wang, Zhongang Qi +1

Despite that convolution neural networks (CNN) have recently demonstrated high-quality reconstruction for video super-resolution (VSR), efficiently training competitive VSR models…

cs.CV20221 cited

RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch Normalization

Xintao Wang, Chao Dong, Ying Shan

This paper explores training efficient VGG-style super-resolution (SR) networks with the structural re-parameterization technique. The general pipeline of re-parameterization is to…

eess.IV20221 cited

VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution

Liangbin Xie. Xintao Wang, Honglun Zhang, Chao Dong +1

Most of the existing video face super-resolution (VFSR) methods are trained and evaluated on VoxCeleb1, which is designed specifically for speaker identification and the frames in…