Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Actionable Interpretability Must Be Defined in Terms of Symmetries
Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini +4
This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpr…
cs.AI2026
Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models
Martino Ciaperoni, Marzio Di Vece, Roberto Pellungrini +3
Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-context lear…