collaborators

5 papers

cs.CV2026

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

Niccolo Avogaro, Nayanika Debnath, Li Mi +6

Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for vision-language models (VLMs). Unstruc…

cs.CV2026

Stitched Value Model for Diffusion Alignment

Hyojun Go, Hyungjin Chung, Prune Truong +8

For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challen…

cs.CV2026

Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

Nicola Farronato, Niccolo Avogaro, Thomas Frick +6

Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed to accurately assess structu…

cs.CV2025

VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation

Niccolo Avogaro, Thomas Frick, Yagmur G. Cinar +12

Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free appr…

cs.CV2025

Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation

Niccolo Avogaro, Thomas Frick, Mattia Rigotti +5

Large Vision-Language Models (VLMs) are increasingly being regarded as foundation models that can be instructed to solve diverse tasks by prompting, without task-specific training.…