425 citations · 571 across the 14 of their papers we have counts for
16 papers · 1 filter
Towards AGI in Computer Vision: Lessons Learned from GPT and Large Language Models
Lingxi Xie, Longhui Wei, Xiaopeng Zhang +4
The AI community has been pursuing algorithms known as artificial general intelligence (AGI) that apply to any kind of real-world problem. Recently, chat systems powered by large l…
Fine-Grained Semantically Aligned Vision-Language Pre-Training
Juncheng Li, Xin He, Longhui Wei +6
Large-scale vision-language pre-training has shown impressive advances in a wide range of downstream tasks. Existing methods mainly model the cross-modal alignment by the similarit…
MVP: Multimodality-guided Visual Pre-training
Longhui Wei, Lingxi Xie, Wengang Zhou +2
Recently, masked image modeling (MIM) has become a promising direction for visual pre-training. In the context of vision transformers, MIM learns effective visual representation by…
Exploring the Diversity and Invariance in Yourself for Visual Pre-Training Task
Longhui Wei, Lingxi Xie, Wengang Zhou +2
Recently, self-supervised learning methods have achieved remarkable success in visual pre-training task. By simply pulling the different augmented views of each image together or o…
What Is Considered Complete for Visual Recognition?
Lingxi Xie, Xiaopeng Zhang, Longhui Wei +2
This is an opinion paper. We hope to deliver a key message that current visual recognition systems are far from complete, i.e., recognizing everything that human can recognize, yet…
Spatiotemporal Transformer for Video-based Person Re-identification
Tianyu Zhang, Longhui Wei, Lingxi Xie +4
Recently, the Transformer module has been transplanted from natural language processing to computer vision. This paper applies the Transformer to video-based person re-identificati…