3 papers
cs.LG2026
Bergson: An Open Source Library for Data Attribution
Lucia Quirke, Louis Jaburi, David Johnston +6
Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging unde…
cs.LG2025
Mechanistic Anomaly Detection for "Quirky" Language Models
David O. Johnston, Arkajyoti Chakraborty, Nora Belrose
As LLMs grow in capability, the task of supervising LLMs becomes more challenging. Supervision failures can occur if LLMs are sensitive to factors that supervisors are unaware of.…
cs.AI2025
Examining Two Hop Reasoning Through Information Content Scaling
David Johnston, Nora Belrose
Prior work has found that transformers have an inconsistent ability to learn to answer latent two-hop questions -- questions of the form "Who is Bob's mother's boss?" We study why…