activity
20222026
most citedRetrieval-Augmented Transformer for Image Captioning

9 citations · 14 across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

21 papers · 1 filter

cs.CV2026

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Davide Caffagni, Alberto Compagnoni, Federico Melis +5

Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an…

cs.CV2026

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

Tobia Poppi, Silvia Cappelletti, Sara Sarto +5

Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interven…

cs.CV2026

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

Evelyn Turri, Davide Bucciarelli, Sara Sarto +2

Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape ima…

cs.CV2026

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models

Gabriele Mattioli, Evelyn Turri, Sara Sarto +3

Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities, and specialized models -- to s…

cs.CV2026

Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering

Marco Morini, Sara Sarto, Marcella Cornia +2

Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visual cues with retrieved textual evidenc…

cs.CV2025

Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models

Davide Caffagni, Sara Sarto, Marcella Cornia +5

Multimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in connecting vision and language, yet their proficiency in fundamental visual reasoning…