collaborators

7 papers

cs.AI2026

Actionable Interpretability Must Be Defined in Terms of Symmetries

Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini +4

This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpr…

cs.LG2026

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga +5

As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, i…

cs.LG2026

Mixture of Concept Bottleneck Experts

Francesco De Santis, Gabriele Ciravegna, Giovanni De Felice +7

Concept Bottleneck Models (CBMs) promote interpretability by grounding predictions in human-understandable concepts. However, existing CBMs typically constrain their task predictor…

cs.AI2026

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

Martino Ciaperoni, Marzio Di Vece, Roberto Pellungrini +3

Large-scale foundation models exhibit \emph{behavioral shifts} when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-conte…

cs.LG2025

If Concept Bottlenecks are the Question, are Foundation Models the Answer?

Nicola Debole, Pietro Barbiero, Francesco Giannini +3

Concept Bottleneck Models (CBMs) are neural networks designed to conjoin high performance with ante-hoc interpretability. CBMs work by first mapping inputs (e.g., images) to high-l…

cs.LG2025

Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts

Andrea Pugnana, Riccardo Massidda, Francesco Giannini +6

Concept Bottleneck Models (CBMs) are machine learning models that improve interpretability by grounding their predictions on human-understandable concepts, allowing for targeted in…