2 papers
cs.LG2026
The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
Eslam Zaher, Maciej Trzaskowski, Quan Nguyen +1
Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of interpretable dictionary atoms,…
cs.LG2026
Counterfactual Explanations on Robust Perceptual Geodesics
Eslam Zaher, Maciej Trzaskowski, Quan Nguyen +1
Latent-space optimization methods for counterfactual explanations - framed as minimal semantic perturbations that change model predictions - inherit the ambiguity of Wachter et al.…