3 papers
cs.LG2026
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
Sophie L. Wang, Phillip Isola, Brian Cheung
How should hidden states generated autoregressively be collapsed into a representation that reflects a language model's internal state? Despite tokens being generated under causal…
cs.LG2025
Enforcing Orderedness to Improve Feature Consistency
Sophie L. Wang, Alex Quach, Nithin Parsan +1
Sparse autoencoders (SAEs) have been widely used for interpretability of neural networks, but their learned features often vary across seeds and hyperparameter settings. We introdu…
cs.CL2025
Words That Make Language Models Perceive
Sophie L. Wang, Phillip Isola, Brian Cheung
Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regular…