3 citations · 11 across the 9 of their papers we have counts for
6 papers · 1 filter
SAM 3: Segment Anything with Concepts
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu +35
We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short…
Pixtral 12B
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +39
We introduce Pixtral-12B, a 12--billion-parameter multimodal language model. Pixtral-12B is trained to understand both natural images and documents, achieving leading performance o…
Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and Benchmarks
Hongjun Wang, Sagar Vaze, Kai Han
Detecting test-time distribution shift has emerged as a key capability for safely deployed machine learning models, with the question being tackled under various guises in recent y…
No Representation Rules Them All in Category Discovery
Sagar Vaze, Andrea Vedaldi, Andrew Zisserman
In this paper we tackle the problem of Generalized Category Discovery (GCD). Specifically, given a dataset with labelled and unlabelled images, the task is to cluster all images in…
GeneCIS: A Benchmark for General Conditional Image Similarity
Sagar Vaze, Nicolas Carion, Ishan Misra
We argue that there are many notions of 'similarity' and that models, like humans, should be able to adapt to these dynamically. This contrasts with most representation learning me…
Optimal Use of Multi-spectral Satellite Data with Convolutional Neural Networks
Sagar Vaze, James Foley, Mohamed Seddiq +2
The analysis of satellite imagery will prove a crucial tool in the pursuit of sustainable development. While Convolutional Neural Networks (CNNs) have made large gains in natural i…