activity
20162024
most citedPaLI: A Jointly-Scaled Multilingual Language-Image Model

196 citations · 714 across the 42 of their papers we have counts for

collaborators
Showing 2022 · cs.CVShow all

5 papers · 2 filters

cs.CV2022★ 1 cited

Rethinking Batch Sample Relationships for Data Representation: A Batch-Graph Transformer based Approach

Xixi Wang, Bo Jiang, Xiao Wang +1

Exploring sample relationships within each mini-batch has shown great potential for learning image representations. Existing works generally adopt the regular Transformer to model…

cs.CV2022★ 196 cited

PaLI: A Jointly-Scaled Multilingual Language-Image Model

Xi Chen, Xiao Wang, Soravit Changpinyo +26

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this…

cs.CV2022★ 3 cited

Few-Shot Learning Meets Transformer: Unified Query-Support Transformers for Few-Shot Classification

Xixi Wang, Xiao Wang, Bo Jiang +1

Few-shot classification which aims to recognize unseen classes using very limited samples has attracted more and more attention. Usually, it is formulated as a metric learning prob…

cs.CV2022★ 66 cited

Simple Open-Vocabulary Object Detection with Vision Transformers

Matthias Minderer, Alexey Gritsenko, Austin Stone +11

Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are…

cs.CV2022★ 2 cited

Tiny Object Tracking: A Large-scale Dataset and A Baseline

Yabin Zhu, Chenglong Li, Yao Liu +4

Tiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in meany vision tasks, such as object detection an…