8 papers
Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B
Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao +5
Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-co…
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez +5
One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answ…
LLMs Process Lists With General Filter Heads
Arnab Sen Sharma, Giordano Rogers, Natalie Shapira +1
We investigate the mechanisms underlying a range of list-processing tasks in LLMs, and we find that LLMs have learned to encode a compact, causal representation of a general filter…
Do explanations generalize across large reasoning models?
Koyena Pal, David Bau, Chandan Singh
Large reasoning models (LRMs) produce a textual chain of thought (CoT) in the process of solving a problem, which serves as a potentially powerful tool to understand the problem by…
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
Hiba Ahsan, Arnab Sen Sharma, Silvio Amir +2
We know from prior work that LLMs encode social biases, and that this manifests in clinical tasks. In this work we adopt tools from mechanistic interpretability to unveil sociodemo…
The Dual-Route Model of Induction
Sheridan Feucht, Eric Todd, Byron Wallace +1
Prior work on in-context copying has shown the existence of induction heads, which attend to and promote individual tokens during copying. In this work we discover a new type of in…