5 papers
Exploring Systems-Thinking Approaches to Loss of Control Risk
Aurelio Carlucci, Sean P. Fillingham, James Walpole +1
Internal deployment of agentic AI systems for coding and research creates a sociotechnical control problem that extends beyond model behaviour. We treat internal-deployment Loss of…
Lessons from External Review of DeepMind's Scheming Inability Safety Case
Stephen Barrett, Francisco Javier Campos Zabala, Sean P. Fillingham +4
Safety cases for frontier AI systems should provide a convincing argument, supported by evidence, that the risk of harm is within an acceptable bound. When developers author their…
STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems
Steve Barrett, Anna Bruvere, Sean P. Fillingham +2
A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The r…
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
Sean P. Fillingham, Andrew Gordon, Peter Lai +3
Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs…
The importance of gas starvation in driving satellite quenching in galaxy groups at
Devontae C. Baxter, Sean P. Fillingham, Alison L. Coil +1
We present results from a Keck/DEIMOS survey to study satellite quenching in group environments at within the Extended Groth Strip (EGS). We target groups in the…