3 papers
cs.LG2025
Distribution-Aware Feature Selection for SAEs
Narmeen Oozeer, Nirmalendu Prakash, Michael Lan +2
Sparse autoencoders (SAEs) decompose neural activations into interpretable features. A widely adopted variant, the TopK SAE, reconstructs each token from its K most active latents.…
cs.LG2025
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
Abir Harrasse, Philip Quirke, Clement Neo +3
Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQ…
cs.AI2025
Activation Space Interventions Can Be Transferred Between Large Language Models
Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash +3
The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representati…