Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Bergson: An Open Source Library for Data Attribution
Lucia Quirke, Louis Jaburi, David Johnston +6
Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging unde…
cs.LG2025
Mechanistic Anomaly Detection for "Quirky" Language Models
David O. Johnston, Arkajyoti Chakraborty, Nora Belrose
As LLMs grow in capability, the task of supervising LLMs becomes more challenging. Supervision failures can occur if LLMs are sensitive to factors that supervisors are unaware of.…