3 papers
cs.AI2024
The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence
Peter Slattery, Alexander K. Saeri, Emily A. C. Grundy +7
Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet researchers, policymakers, and technology c…
cs.LG2024
Challenges in Mechanistically Interpreting Model Representations
Satvik Golechha, James Dao
Mechanistic interpretability (MI) aims to understand AI models by reverse-engineering the exact algorithms neural networks learn. Most works in MI so far have studied behaviors and…
cs.LG2023
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
Jett Janiak, Can Rager, James Dao +1
Prior work suggests that language models manage the limited bandwidth of the residual stream through a "memory management" mechanism, where certain attention heads and MLP layers c…