85 citations · 174 across the 23 of their papers we have counts for
27 papers
Darwinian Model Upgrades: Model Evolving with Selective Compatibility
Binjie Zhang, Shupeng Su, Yixiao Ge +5
The traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensiv…
MonoNeuralFusion: Online Monocular Neural 3D Reconstruction with Geometric Priors
Zi-Xin Zou, Shi-Sheng Huang, Yan-Pei Cao +3
High-fidelity 3D scene reconstruction from monocular videos continues to be challenging, especially for complete and fine-grained geometry reconstruction. The previous 3D reconstru…
Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection
Yuxin Fang, Shusheng Yang, Shijie Wang +3
We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…
Accelerating the Training of Video Super-Resolution Models
Lijian Lin, Xintao Wang, Zhongang Qi +1
Despite that convolution neural networks (CNN) have recently demonstrated high-quality reconstruction for video super-resolution (VSR), efficiently training competitive VSR models…
RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch Normalization
Xintao Wang, Chao Dong, Ying Shan
This paper explores training efficient VGG-style super-resolution (SR) networks with the structural re-parameterization technique. The general pipeline of re-parameterization is to…
VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution
Liangbin Xie. Xintao Wang, Honglun Zhang, Chao Dong +1
Most of the existing video face super-resolution (VFSR) methods are trained and evaluated on VoxCeleb1, which is designed specifically for speaker identification and the frames in…