activity
20182023
most citedXCiT: Cross-Covariance Image Transformers

234 citations · 636 across the 14 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

5 papers · 2 filters

cs.CV2023★ 26 cited

Self-Supervised Learning for Endoscopic Video Analysis

Roy Hirsch, Mathilde Caron, Regev Cohen +6

Self-supervised learning (SSL) has led to important breakthroughs in computer vision by allowing learning from large amounts of unlabeled data. As such, it might have a pivotal rol…

cs.CV2023★ 16 cited

Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution

Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12

The ubiquitous and demonstrably suboptimal choice of resizing images to a fixed resolution before processing them with computer vision models has not yet been successfully challeng…

cs.CV2023★ 4 cited

Retrieval-Enhanced Contrastive Vision-Text Models

Ahmet Iscen, Mathilde Caron, Alireza Fathi +1

Contrastive image-text models such as CLIP form the building blocks of many state-of-the-art systems. While they excel at recognizing common generic concepts, they still struggle o…

cs.CV2023

Verbs in Action: Improving verb understanding in video-language models

Liliane Momeni, Mathilde Caron, Arsha Nagrani +2

Understanding verbs is crucial to modelling how people and objects interact with each other and the environment through space and time. Recently, state-of-the-art video-language mo…

cs.CV2023★ 118 cited

Scaling Vision Transformers to 22 Billion Parameters

Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…