28 papers
Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik
Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists…
KnowsTFM: Knowledge-Informed Fine-Tuning of Small Tabular Foundation Models
Boshko Koloski, Xiangjian Jiang, Senja Pollak +3
Tabular foundation models have advanced deep learning for tabular data by delivering strong default performance across many small and medium tasks. Yet in niche domains, where data…
Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
Tiansi Dong, Mateja Jamnik, Pietro Liò
By promoting vectors to spheres and enabling explicit model construction, neural networks can perform symbolic-level syllogistic reasoning without training data. We identify two fu…
Actionable Interpretability Must Be Defined in Terms of Symmetries
Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini +4
This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpr…
The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics
Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga +5
As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, i…
CB-SLICE: Concept-Based Interpretable Error Slice Discovery
Yael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud +1
Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Identifying these groups and the…