Publications (16)
Large-Scale Attribute-Object Compositions
Filip Radenovic, Animesh Sinha, Albert Gordo +2
We study the problem of learning how to predict attribute-object compositions from images, and its generalization to unseen compositions missing from the training data. To the best…
Repeatability Is Not Enough: Learning Affine Regions via Discriminability
Dmytro Mishkin, Filip Radenovic, Jiri Matas
A method for learning local affine-covariant regions is presented. We show that maximizing geometric repeatability does not lead to local regions, a.k.a features,that are reliably…
Filtering, Distillation, and Hard Negatives for Vision-Language Pre-Training
Filip Radenovic, Abhimanyu Dubey, Abhishek Kadian +6
Vision-language models trained with contrastive learning on large-scale noisy data are becoming increasingly popular for zero-shot recognition problems. In this paper we improve th…
Scalable Interpretability via Polynomials
Abhimanyu Dubey, Filip Radenovic, Dhruv Mahajan
Generalized Additive Models (GAMs) have quickly become the leading choice for inherently-interpretable machine learning. However, unlike uninterpretable methods such as DNNs, they…
Neural Basis Models for Interpretability
Filip Radenovic, Abhimanyu Dubey, Dhruv Mahajan
Due to the widespread use of complex machine learning models in real-world applications, it is becoming critical to explain model predictions. However, these models are typically b…
Attention-Based Query Expansion Learning
Albert Gordo, Filip Radenovic, Tamara Berg
Query expansion is a technique widely used in image search consisting in combining highly ranked images from an original query into an expanded query that is then reissued, general…
Context Diffusion: In-Context Aware Image Generation
Ivona Najdenkoska, Animesh Sinha, Abhimanyu Dubey +3
We propose Context Diffusion, a diffusion-based framework that enables image generation models to learn from visual examples presented in context. Recent work tackles such in-conte…
COLA: A Benchmark for Compositional Text-to-image Retrieval
Arijit Ray, Filip Radenovic, Abhimanyu Dubey +3
Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combini…
Multiple Measurements and Joint Dimensionality Reduction for Large Scale Image Search with Short Vectors - Extended Version
Filip Radenovic, Herve Jegou, Ondrej Chum
This paper addresses the construction of a short-vector (128D) image representation for large-scale image and particular object retrieval. In particular, the method of joint dimens…
The 2021 Image Similarity Dataset and Challenge
Matthijs Douze, Giorgos Tolias, Ed Pizzi +9
This paper introduces a new benchmark for large-scale image similarity detection. This benchmark is used for the Image Similarity Challenge at NeurIPS'21 (ISC2021). The goal is to…
Camera Elevation Estimation from a Single Mountain Landscape Photograph
Martin Cadik, Jan Vasicek, Michal Hradis +2
This work addresses the problem of camera elevation estimation from a single photograph in an outdoor environment. We introduce a new benchmark dataset of one-hundred thousand imag…
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models th…
Making Heads or Tails: Towards Semantically Consistent Visual Counterfactuals
Simon Vandenhende, Dhruv Mahajan, Filip Radenovic +1
A visual counterfactual explanation replaces image regions in a query image with regions from a distractor image such that the system's decision on the transformed image changes to…
Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
Xiaoliang Dai, Ji Hou, Chih-Yao Ma +23
Training text-to-image models with web scale image-text pairs enables the generation of a wide range of visual concepts from text. However, these pre-trained models often face chal…
Targeted Mismatch Adversarial Attack: Query with a Flower to Retrieve the Tower
Giorgos Tolias, Filip Radenovic, OndÅej Chum
Access to online visual search engines implies sharing of private user content - the query images. We introduce the concept of targeted mismatch attack for deep learning based retr…
Working hard to know your neighbor's margins: Local descriptor learning loss
Anastasiya Mishchuk, Dmytro Mishkin, Filip Radenovic +1
We introduce a novel loss for learning local feature descriptors which is inspired by the Lowe's matching criterion for SIFT. We show that the proposed loss that maximizes the dist…