2 papers
cs.LG2026
Building Better Activation Oracles
Jan Bauer, Celeste De Schamphelaere, Adam Karvonen +2
Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallucinations and vagueness. Addit…
cs.LG2024
Universal Response and Emergence of Induction in LLMs
Niclas Luick
While induction is considered a key mechanism for in-context learning in LLMs, understanding its precise circuit decomposition beyond toy models remains elusive. Here, we study the…