7 papers
Actionable Interpretability Must Be Defined in Terms of Symmetries
Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini +4
This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpr…
The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics
Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga +5
As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, i…
Mixture of Concept Bottleneck Experts
Francesco De Santis, Gabriele Ciravegna, Giovanni De Felice +7
Concept Bottleneck Models (CBMs) promote interpretability by grounding predictions in human-understandable concepts. However, existing CBMs typically constrain their task predictor…
Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models
Martino Ciaperoni, Marzio Di Vece, Roberto Pellungrini +3
Large-scale foundation models exhibit \emph{behavioral shifts} when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-conte…
If Concept Bottlenecks are the Question, are Foundation Models the Answer?
Nicola Debole, Pietro Barbiero, Francesco Giannini +3
Concept Bottleneck Models (CBMs) are neural networks designed to conjoin high performance with ante-hoc interpretability. CBMs work by first mapping inputs (e.g., images) to high-l…
Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts
Andrea Pugnana, Riccardo Massidda, Francesco Giannini +6
Concept Bottleneck Models (CBMs) are machine learning models that improve interpretability by grounding their predictions on human-understandable concepts, allowing for targeted in…