23 citations · 56 across the 9 of their papers we have counts for
11 papers
Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
Hao Li, Jinguo Zhu, Xiaohu Jiang +8
Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to elimi…
Darwinian Model Upgrades: Model Evolving with Selective Compatibility
Binjie Zhang, Shupeng Su, Yixiao Ge +5
The traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensiv…
DeViT: Deformed Vision Transformers in Video Inpainting
Jiayin Cai, Changlin Li, Xin Tao +2
This paper proposes a novel video inpainting method. We make three main contributions: First, we extended previous Transformers with patch alignment by introducing Deformed Patch-b…
ViTKD: Practical Guidelines for ViT feature knowledge distillation
Zhendong Yang, Zhe Li, Ailing Zeng +3
Knowledge Distillation (KD) for Convolutional Neural Network (CNN) is extensively studied as a way to boost the performance of a small model. Recently, Vision Transformer (ViT) has…
Privacy-Preserving Model Upgrades with Bidirectional Compatible Training in Image Retrieval
Shupeng Su, Binjie Zhang, Yixiao Ge +4
The task of privacy-preserving model upgrades in image retrieval desires to reap the benefits of rapidly evolving new models without accessing the raw gallery images. A pioneering…
Towards Universal Backward-Compatible Representation Learning
Binjie Zhang, Yixiao Ge, Yantao Shen +6
Conventional model upgrades for visual search systems require offline refresh of gallery features by feeding gallery images into new models (dubbed as "backfill"), which is time-co…