collaborators
Showing cs.LGShow all

17 papers · 1 filter

cs.LG2026

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik

Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists…

cs.LG2026

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga +5

As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, i…

cs.LG2026

In Defense of Information Leakage in Concept-based Models

Mateo Espinosa Zarlenga

Concept-based models (CMs), deep neural networks that ground their predictions on representations aligned with human-understandable concepts (e.g., "round", "stripes", etc.), have…

cs.LG2026

Mixture of Concept Bottleneck Experts

Francesco De Santis, Gabriele Ciravegna, Giovanni De Felice +7

Concept Bottleneck Models (CBMs) promote interpretability by grounding predictions in human-understandable concepts. However, existing CBMs typically constrain their task predictor…

cs.LG2026

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

Yael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud +1

Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Identifying these groups and the…

cs.LG2026

Digging Deeper: Learning Multi-Level Concept Hierarchies

Oscar Hill, Mateo Espinosa Zarlenga, Mateja Jamnik

Although concept-based models promise interpretability by explaining predictions with human-understandable concepts, they typically rely on exhaustive annotations and treat concept…