2 papers
cs.CL2026
Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models
Quintin Pope, Ajay Hayagreeve Balaji, Jacques Thibodeau +1
We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model and an intervention…
cs.LG2024
Neural Networks Learn Statistics of Increasing Complexity
Nora Belrose, Quintin Pope, Lucia Quirke +2
The distributional simplicity bias (DSB) posits that neural networks learn low-order moments of the data distribution first, before moving on to higher-order correlations. In this…