activity
20172025
most citedTraining data-efficient image transformers & distillation through attention

150 citations · 439 across the 32 of their papers we have counts for

collaborators
Showing cs.CVShow all

55 papers · 1 filter

cs.CV2025

JAFAR: Jack up Any Feature at Any Resolution

Paul Couairon, Loick Chambon, Louis Serrano +3

Foundation Vision Encoders have become essential for a wide range of dense vision tasks. However, their low-resolution spatial feature outputs necessitate feature upsampling to pro…

cs.CV2023

Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning

Mustafa Shukor, Alexandre Rame, Corentin Dancette +1

Following the success of Large Language Models (LLMs), Large Multimodal Models (LMMs), such as the Flamingo model and its subsequent competitors, have started to emerge as natural…

cs.CV2023

UnIVAL: Unified Model for Image, Video, Audio and Language Tasks

Mustafa Shukor, Corentin Dancette, Alexandre Rame +1

Large Language Models (LLMs) have made the ambitious quest for generalist agents significantly far from being a fantasy. A key hurdle for building such general models is the divers…

cs.CV2023

MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments

Spyros Gidaris, Andrei Bursuc, Oriane Simeoni +4

Self-supervised learning can be used for mitigating the greedy needs of Vision Transformer networks for very large fully-annotated datasets. Different classes of self-supervised le…

cs.CV2023

Zero-shot spatial layout conditioning for text-to-image diffusion models

Guillaume Couairon, Marlène Careil, Matthieu Cord +2

Large-scale text-to-image diffusion models have significantly improved the state of the art in generative image modelling and allow for an intuitive and powerful user interface to…

cs.CV20231 cited

Improving Selective Visual Question Answering by Learning from Your Peers

Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary +5

Despite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains underexplored. Recent work has shown that VQA models, out-of-the-…