3 papers
cs.LG2025
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…
cs.CL2025
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
Amir Zur, Atticus Geiger, Ekdeep Singh Lubana +1
When a language model generates text, the selection of individual tokens might lead it down very different reasoning paths, making uncertainty difficult to quantify. In this work,…
cs.CV2025
Chain of Time: In-Context Physical Simulation with Image Generation Models
YingQiao Wang, Eric Bigelow, Boyi Li +1
We propose a novel cognitively-inspired method to improve and interpret physical simulation in vision-language models. Our ``Chain of Time" method involves generating a series of i…