papers

Publications (16)

cs.CV2021

Large-Scale Attribute-Object Compositions

Filip Radenovic, Animesh Sinha, Albert Gordo +2

We study the problem of learning how to predict attribute-object compositions from images, and its generalization to unseen compositions missing from the training data. To the best…

cs.CV2018

Repeatability Is Not Enough: Learning Affine Regions via Discriminability

Dmytro Mishkin, Filip Radenovic, Jiri Matas

A method for learning local affine-covariant regions is presented. We show that maximizing geometric repeatability does not lead to local regions, a.k.a features,that are reliably…

cs.CV2023

Filtering, Distillation, and Hard Negatives for Vision-Language Pre-Training

Filip Radenovic, Abhimanyu Dubey, Abhishek Kadian +6

Vision-language models trained with contrastive learning on large-scale noisy data are becoming increasingly popular for zero-shot recognition problems. In this paper we improve th…

cs.LG2022

Scalable Interpretability via Polynomials

Abhimanyu Dubey, Filip Radenovic, Dhruv Mahajan

Generalized Additive Models (GAMs) have quickly become the leading choice for inherently-interpretable machine learning. However, unlike uninterpretable methods such as DNNs, they…

cs.LG2022

Neural Basis Models for Interpretability

Filip Radenovic, Abhimanyu Dubey, Dhruv Mahajan

Due to the widespread use of complex machine learning models in real-world applications, it is becoming critical to explain model predictions. However, these models are typically b…

cs.CV2020

Attention-Based Query Expansion Learning

Albert Gordo, Filip Radenovic, Tamara Berg

Query expansion is a technique widely used in image search consisting in combining highly ranked images from an original query into an expanded query that is then reissued, general…

cs.CV2025

Context Diffusion: In-Context Aware Image Generation

Ivona Najdenkoska, Animesh Sinha, Abhimanyu Dubey +3

We propose Context Diffusion, a diffusion-based framework that enables image generation models to learn from visual examples presented in context. Recent work tackles such in-conte…

cs.CV2023

COLA: A Benchmark for Compositional Text-to-image Retrieval

Arijit Ray, Filip Radenovic, Abhimanyu Dubey +3

Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combini…

cs.CV2015

Multiple Measurements and Joint Dimensionality Reduction for Large Scale Image Search with Short Vectors - Extended Version

Filip Radenovic, Herve Jegou, Ondrej Chum

This paper addresses the construction of a short-vector (128D) image representation for large-scale image and particular object retrieval. In particular, the method of joint dimens…

cs.CV2022

The 2021 Image Similarity Dataset and Challenge

Matthijs Douze, Giorgos Tolias, Ed Pizzi +9

This paper introduces a new benchmark for large-scale image similarity detection. This benchmark is used for the Image Similarity Challenge at NeurIPS'21 (ISC2021). The goal is to…

cs.CV2016

Camera Elevation Estimation from a Single Mountain Landscape Photograph

Martin Cadik, Jan Vasicek, Michal Hradis +2

This work addresses the problem of camera elevation estimation from a single photograph in an outdoor environment. We introduce a new benchmark dataset of one-hundred thousand imag…

cs.AI2024

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +556

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models th…

cs.CV2022

Making Heads or Tails: Towards Semantically Consistent Visual Counterfactuals

Simon Vandenhende, Dhruv Mahajan, Filip Radenovic +1

A visual counterfactual explanation replaces image regions in a query image with regions from a distractor image such that the system's decision on the transformed image changes to…

cs.CV2023

Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Xiaoliang Dai, Ji Hou, Chih-Yao Ma +23

Training text-to-image models with web scale image-text pairs enables the generation of a wide range of visual concepts from text. However, these pre-trained models often face chal…

cs.CV2019

Targeted Mismatch Adversarial Attack: Query with a Flower to Retrieve the Tower

Giorgos Tolias, Filip Radenovic, Ondřej Chum

Access to online visual search engines implies sharing of private user content - the query images. We introduce the concept of targeted mismatch attack for deep learning based retr…

cs.CV2018

Working hard to know your neighbor's margins: Local descriptor learning loss

Anastasiya Mishchuk, Dmytro Mishkin, Filip Radenovic +1

We introduce a novel loss for learning local feature descriptors which is inspired by the Lowe's matching criterion for SIFT. We show that the proposed loss that maximizes the dist…