collaborators

6 papers

cs.CV2026

Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

Mateo Diaz-Bone, Daniel Caraballo, Florian Scheidegger +11

Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these methods face challenges when…

cs.CV2026

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning

Brown Ebouky, Gabriele Carrino, Niccolo Avogaro +3

Human visual reasoning is governed by active vision, a process where metacognitive control drives top-down goal-directed attention, dynamically routing foveal focus toward task-rel…

cs.CV2026

Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training

Brown Ebouky, Ajad Chhatkuli, Cristiano Malossi +3

Self-supervised learning (SSL) has emerged as a central paradigm for training foundation models by leveraging large-scale unlabeled datasets, often producing representations with s…

cs.CL2025

Eliciting Reasoning in Language Models with Cognitive Tools

Brown Ebouky, Andrea Bartezzaghi, Mattia Rigotti

The recent advent of reasoning models like OpenAI's o1 was met with excited speculation by the AI community about the mechanisms underlying these capabilities in closed models, fol…

cs.CV2025

Advanced Layout Analysis Models for Docling

Nikolaos Livathinos, Christoph Auer, Ahmed Nassar +16

This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object…

cs.CV2025

VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation

Niccolo Avogaro, Thomas Frick, Yagmur G. Cinar +12

Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free appr…