3 citations · 3 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability
Dibyanayan Bandyopadhyay, Asif Ekbal
Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be…
cs.LG2026
Sparse Semantic Dimension as a Generalization Certificate for LLMs
Dibyanayan Bandyopadhyay, Asif Ekbal
Standard statistical learning theory predicts that Large Language Models (LLMs) should overfit because their parameter counts vastly exceed the number of training tokens. Yet, in p…