9 papers
OASIC: Occlusion-Agnostic and Severity-Informed Classification
Kay Gijzen, Gertjan J. Burghouts, Daniël M. Pelt
Severe occlusions of objects pose a major challenge for computer vision. We show that two root causes are (1) the loss of visible information and (2) the distracting patterns cause…
Better Language Models Exhibit Higher Visual Alignment
Jona Ruthardt, Gertjan J. Burghouts, Serge Belongie +1
How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of vario…
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
Emanuele Mezzi, Gertjan Burghouts, Maarten Kruithof
Text-to-image retrieval in remote sensing (RS) has advanced rapidly with the rise of large vision-language models (LVLMs) tailored for aerial and satellite imagery, culminating in…
Occlusion Robustness of CLIP for Military Vehicle Classification
Jan Erik van Woerden, Gertjan Burghouts, Lotte Nijskens +4
Vision-language models (VLMs) like CLIP enable zero-shot classification by aligning images and text in a shared embedding space, offering advantages for defense applications with s…
Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
Frank Ruis, Gertjan Burghouts, Hugo Kuijf
Recent progress in large pre-trained vision language models (VLMs) has reached state-of-the-art performance on several object detection benchmarks and boasts strong zero-shot capab…
Near, far: Patch-ordering enhances vision foundation models' scene understanding
Valentinos Pariza, Mohammadreza Salehi, Gertjan Burghouts +2
We introduce NeCo: Patch Neighbor Consistency, a novel self-supervised training loss that enforces patch-level nearest neighbor consistency across a student and teacher model. Comp…