11 papers · 1 filter
MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
Oskar Kristoffersen, Alba Reinders Sánchez, Morten Rieger Hannemose +2
Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textua…
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
Marco Schouten, Ioannis Siglidis, Serge Belongie +1
We propose a method to learn explicit, class-conditioned spatial priors for object placement in natural scenes by distilling the implicit placement knowledge encoded in text-condit…
Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels
Juan Miguel Valverde, Dim P. Papadopoulos, Rasmus Larsen +1
Standard deep learning models for image segmentation cannot guarantee topology accuracy, failing to preserve the correct number of connected components or structures. This, in turn…
Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
Kaixuan Lu, Mehmet Onurcan Kaya, Dim P. Papadopoulos
Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervis…
Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
Erik Riise, Mehmet Onurcan Kaya, Dim P. Papadopoulos
While inference-time scaling through search has revolutionized Large Language Models, translating these gains to image generation has proven difficult. Recent attempts to apply sea…
AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
Kaixuan Lu, Mehmet Onurcan Kaya, Dim P. Papadopoulos
Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervis…