6 papers
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
Niccolo Avogaro, Nayanika Debnath, Li Mi +6
Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for vision-language models (VLMs). Unstruc…
Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision
Mateo Diaz-Bone, Daniel Caraballo, Florian Scheidegger +11
Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these methods face challenges when…
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
Nicola Farronato, Niccolo Avogaro, Thomas Frick +6
Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed to accurately assess structu…
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
Brown Ebouky, Gabriele Carrino, Niccolo Avogaro +3
Human visual reasoning is governed by active vision, a process where metacognitive control drives top-down goal-directed attention, dynamically routing foveal focus toward task-rel…
VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation
Niccolo Avogaro, Thomas Frick, Yagmur G. Cinar +12
Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free appr…
Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation
Niccolo Avogaro, Thomas Frick, Mattia Rigotti +5
Large Vision-Language Models (VLMs) are increasingly being regarded as foundation models that can be instructed to solve diverse tasks by prompting, without task-specific training.…