2 papers
cs.AI2026
Supervised sparse auto-encoders for interpretable and compositional representations
Ouns El Harzli, Hugo Wallner, Yoonsoo Nam +1
Sparse auto-encoders (SAEs) have re-emerged as a prominent method for mechanistic interpretability, yet they face two significant challenges: the non-smoothness of the penalt…
cs.LG2026
Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks
Ouns El Harzli, Yoonsoo Nam, Ilja Kuzborskij +2
Algorithmic stability is a classical framework for analyzing the generalization error of learning algorithms. It predicts that an algorithm has small generalization error if it is…