3 papers
cs.LG2026
Mechanistic Interpretability of Antibody Language Models Using SAEs
Rebonto Haque, Oliver M. Turnbull, Anisha Parsan +4
Sparse autoencoders (SAEs) are a mechanistic interpretability technique that have been used to provide insight into learned concepts within large protein language models. Here, we…
cs.LG2025
Enforcing Orderedness to Improve Feature Consistency
Sophie L. Wang, Alex Quach, Nithin Parsan +1
Sparse autoencoders (SAEs) have been widely used for interpretability of neural networks, but their learned features often vary across seeds and hyperparameter settings. We introdu…
q-bio.BM2025
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
Nithin Parsan, David J. Yang, John J. Yang
Protein language models have revolutionized structure prediction, but their nonlinear nature obscures how sequence representations inform structure prediction. While sparse autoenc…