Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Challenges in Mechanistically Interpreting Model Representations
Satvik Golechha, James Dao
Mechanistic interpretability (MI) aims to understand AI models by reverse-engineering the exact algorithms neural networks learn. Most works in MI so far have studied behaviors and…
cs.LG2023
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
Jett Janiak, Can Rager, James Dao +1
Prior work suggests that language models manage the limited bandwidth of the residual stream through a "memory management" mechanism, where certain attention heads and MLP layers c…