8 papers
The Capability Frontier: Benchmarks Miss 82% of Model Performance
Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8
Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data…
Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
Philip Quirke
Mechanistic studies often assign a component a role when removing it damages a behavior, its activation linearly encodes task information, and restoring that activation repairs the…
Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering
Narmeen Oozeer, Shivam Raval, Philip Quirke +4
Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation to nonlinear methods such as…
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing
Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi +4
While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As…
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
Abir Harrasse, Philip Quirke, Clement Neo +3
Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQ…
Position: Require Frontier AI Labs To Release Small "Analog" Models
Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer +1
Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovati…