Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta +4
In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations…
cs.AI2025
Logic Explanation of AI Classifiers by Categorical Explaining Functors
Stefano Fioravanti, Francesco Giannini, Paolo Frazzetto +2
The most common methods in explainable artificial intelligence are post-hoc techniques which identify the most relevant features used by pretrained opaque models. Some of the most…