Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Improving Robustness In Sparse Autoencoders via Masked Regularization
Vivek Narayanaswamy, Kowshik Thopalli, Bhavya Kailkhura +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability to project LLM activations onto sparse latent spaces. However, sparsity alone is an imperfect proxy for i…
cs.LG2024
On the Use of Anchoring for Training Vision Models
Vivek Narayanaswamy, Kowshik Thopalli, Rushil Anirudh +3
Anchoring is a recent, architecture-agnostic principle for training deep neural networks that has been shown to significantly improve uncertainty estimation, calibration, and extra…