activity
20212023
most citedSEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

52 citations · 259 across the 52 of their papers we have counts for

collaborators

7 papers

cs.CV20224 cited

Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection

Yuxin Fang, Shusheng Yang, Shijie Wang +3

We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…

cs.CV2022

Accelerating the Training of Video Super-Resolution Models

Lijian Lin, Xintao Wang, Zhongang Qi +1

Despite that convolution neural networks (CNN) have recently demonstrated high-quality reconstruction for video super-resolution (VSR), efficiently training competitive VSR models…

cs.CV20221 cited

RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch Normalization

Xintao Wang, Chao Dong, Ying Shan

This paper explores training efficient VGG-style super-resolution (SR) networks with the structural re-parameterization technique. The general pipeline of re-parameterization is to…

eess.IV20221 cited

VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution

Liangbin Xie. Xintao Wang, Honglun Zhang, Chao Dong +1

Most of the existing video face super-resolution (VFSR) methods are trained and evaluated on VoxCeleb1, which is designed specifically for speaker identification and the frames in…

cs.CV20223 cited

Privacy-Preserving Model Upgrades with Bidirectional Compatible Training in Image Retrieval

Shupeng Su, Binjie Zhang, Yixiao Ge +4

The task of privacy-preserving model upgrades in image retrieval desires to reap the benefits of rapidly evolving new models without accessing the raw gallery images. A pioneering…

cs.CV2022

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

Yuying Ge, Yixiao Ge, Xihui Liu +5

Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast gl…