2 papers
cs.CL2026
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible re…
cs.CL2026
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
Kevin Der, Harish Kamath, Ben Thompson
Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models. However, standard SAE architectures operate on individual token activ…