106 citations · 152 across the 9 of their papers we have counts for
13 papers · 1 filter
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
Loick Chambon, Paul Couairon, Eloi Zablocki +3
Vision Foundation Models (VFMs) extract spatially downsampled representations, posing challenges for pixel-level tasks. Existing upsampling approaches face a fundamental trade-off:…
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
Sophia Sirko-Galouchenko, Spyros Gidaris, Antonin Vobecky +2
We introduce DIP, a novel unsupervised post-training method designed to enhance dense image representations in large-scale pretrained vision encoders for in-context scene understan…
JAFAR: Jack up Any Feature at Any Resolution
Paul Couairon, Loick Chambon, Louis Serrano +3
Foundation Vision Encoders have become essential for a wide range of dense vision tasks. However, their low-resolution spatial feature outputs necessitate feature upsampling to pro…
Memory transformers for full context and high-resolution 3D Medical Segmentation
Loic Themyr, Clément Rambour, Nicolas Thome +2
Transformer models achieve state-of-the-art results for image segmentation. However, achieving long-range attention, necessary to capture global context, with high-resolution 3D im…
Swapping Semantic Contents for Mixing Images
Rémy Sun, Clément Masson, Gilles Hénaff +2
Deep architecture have proven capable of solving many tasks provided a sufficient amount of labeled data. In fact, the amount of available labeled data has become the principal bot…
Confidence Estimation via Auxiliary Models
Charles Corbière, Nicolas Thome, Antoine Saporta +3
Reliably quantifying the confidence of deep neural classifiers is a challenging yet fundamental requirement for deploying such models in safety-critical applications. In this paper…