5 papers
JAFAR: Jack up Any Feature at Any Resolution
Paul Couairon, Loick Chambon, Louis Serrano +3
Foundation Vision Encoders have become essential for a wide range of dense vision tasks. However, their low-resolution spatial feature outputs necessitate feature upsampling to pro…
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
Loick Chambon, Paul Couairon, Eloi Zablocki +3
Vision Foundation Models (VFMs) extract spatially downsampled representations, posing challenges for pixel-level tasks. Existing upsampling approaches face a fundamental trade-off:…
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
Paul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard +2
Foundation models have emerged as powerful tools across various domains including language, vision, and multimodal tasks. While prior works have addressed unsupervised image segmen…
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
Sophia Sirko-Galouchenko, Spyros Gidaris, Antonin Vobecky +2
We introduce DIP, a novel unsupervised post-training method designed to enhance dense image representations in large-scale pretrained vision encoders for in-context scene understan…
Hierarchical Clustering in Mean-Field Coupled Stuart-Landau Oscillators
Nicolas Thomé, Matthias Wolfrum, Katharina Krischer
Clustered solutions in oscillator networks provide an important insight into how a system might diversify from a synchronous solution into spatiotemporal complex solutions. They ca…