activity
20232026
most citedA Configurable Library for Generating and Manipulating Maze Datasets

2 citations · 2 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CL2026

Capability Provenance in Language Models: A Case Study in Social Reasoning

Glenn Matlin, Chandreyi Chakraborty, Saehee Eom +8

We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning i…

cs.LG2026

Bergson: An Open Source Library for Data Attribution

Lucia Quirke, Louis Jaburi, David Johnston +6

Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging unde…

cs.LG2025

Binary Sparse Coding for Interpretability

Lucia Quirke, Stepan Shabalin, Nora Belrose

Sparse autoencoders (SAEs) are used to decompose neural network activations into sparsely activating features, but many SAE features are only interpretable at high activation stren…

cs.LG2025

Slowing Learning by Erasing Simple Features

Lucia Quirke, Nora Belrose

Prior work suggests that neural networks tend to learn low-order moments of the data distribution first, before moving on to higher-order correlations. In this work, we derive a no…

cs.LG2024

Neural Networks Learn Statistics of Increasing Complexity

Nora Belrose, Quintin Pope, Lucia Quirke +2

The distributional simplicity bias (DSB) posits that neural networks learn low-order moments of the data distribution first, before moving on to higher-order correlations. In this…

cs.LG2023

Structured World Representations in Maze-Solving Transformers

Michael Igorevich Ivanitskiy, Alex F. Spies, Tilman Räuker +9

Transformer models underpin many recent advances in practical machine learning applications, yet understanding their internal behavior continues to elude researchers. Given the siz…