2 papers
cs.LG2026
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
Sai V R Chereddy
The paper explores how video models trained for classification tasks represent nuanced, hidden semantic information that may not affect the final outcome, a key challenge for Trust…
cs.CL2025
ParaScopes: What do Language Models Activations Encode About Future Text?
Nicky Pochinkov, Yulia Volkova, Anna Vasileva +1
Interpretability studies in language models often investigate forward-looking representations of activations. However, as language models become capable of doing ever longer time h…