6 papers
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta +4
In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations…
Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling
David Méndez, Roberto Confalonieri, Natalia DÃaz RodrÃguez
Vision-Language Models (VLMs) excel at tasks like zero-shot classification and cross-modal retrieval by mapping images and text to a shared space, but this requires expensive end-t…
Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
Matteo Gioele Collu, Umberto Salviati, Roberto Confalonieri +2
Large Language Models (LLMs) are increasingly being integrated into the scientific peer-review process, raising new questions about their reliability and resilience to manipulation…
Iterative In-Context Learning to Enhance LLMs Abstract Reasoning: The Case-Study of Algebraic Tasks
Stefano Fioravanti, Matteo Zavatteri, Roberto Confalonieri +4
LLMs face significant challenges in systematic generalization, particularly when dealing with reasoning tasks requiring compositional rules and handling out-of-distribution example…
CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs
David Méndez, Gianpaolo Bontempo, Elisa Ficarra +2
Deep vision models often rely on biases learned from spurious correlations in datasets. To identify these biases, methods that interpret high-level, human-understandable concepts a…
Logic Explanation of AI Classifiers by Categorical Explaining Functors
Stefano Fioravanti, Francesco Giannini, Paolo Frazzetto +2
The most common methods in explainable artificial intelligence are post-hoc techniques which identify the most relevant features used by pretrained opaque models. Some of the most…