282 citations · 356 across the 70 of their papers we have counts for
14 papers · 2 filters
DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering
Luca De Grandis, Silvia Cappelletti, William Raccagni +3
Answer grounding in document visual question answering remains an open challenge: most benchmarks lack grounding annotations or provide limited-quality labels, while constructing g…
SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution
Federico Putamorsi, Leonardo Zini, Marcella Cornia +1
Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose internal activations can be dominated by a small number of massive channe…
A Scalable Vector Graphics Latent Space
Leonardo Zini, Elia Frigieri, Lorenzo Baraldi
Scalable Vector Graphics are a fundamental medium for resolution-independent visual content, yet the deep learning community lacks a continuous, dense, and invertible latent space…
Mind the Heads: Topological Representation Alignment for Multimodal LLMs
Davide Caffagni, Alberto Compagnoni, Federico Melis +5
Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an…
SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation
Zeno Testa, Antonino Furnari, Lorenzo Baraldi +1
Sign Language Translation (SLT) is typically evaluated with surface-form metrics such as BLEU and ROUGE, which reward lexical overlap but do not directly measure whether a translat…
Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation
Tobia Poppi, Silvia Cappelletti, Sara Sarto +5
Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interven…