activity
20242026
collaborators

5 papers

cs.CY2026

Exploring Systems-Thinking Approaches to Loss of Control Risk

Aurelio Carlucci, Sean P. Fillingham, James Walpole +1

Internal deployment of agentic AI systems for coding and research creates a sociotechnical control problem that extends beyond model behaviour. We treat internal-deployment Loss of…

cs.CY2026

Lessons from External Review of DeepMind's Scheming Inability Safety Case

Stephen Barrett, Francisco Javier Campos Zabala, Sean P. Fillingham +4

Safety cases for frontier AI systems should provide a convincing argument, supported by evidence, that the risk of harm is within an acceptable bound. When developers author their…

cs.CY2026

STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems

Steve Barrett, Anna Bruvere, Sean P. Fillingham +2

A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The r…

cs.LG2025

SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs

Sean P. Fillingham, Andrew Gordon, Peter Lai +3

Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs…

astro-ph.GA2024

The importance of gas starvation in driving satellite quenching in galaxy groups at

Devontae C. Baxter, Sean P. Fillingham, Alison L. Coil +1

We present results from a Keck/DEIMOS survey to study satellite quenching in group environments at within the Extended Groth Strip (EGS). We target groups in the…