Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
Exploring Systems-Thinking Approaches to Loss of Control Risk
Aurelio Carlucci, Sean P. Fillingham, James Walpole +1
Internal deployment of agentic AI systems for coding and research creates a sociotechnical control problem that extends beyond model behaviour. We treat internal-deployment Loss of…
cs.CY2026
Lessons from External Review of DeepMind's Scheming Inability Safety Case
Stephen Barrett, Francisco Javier Campos Zabala, Sean P. Fillingham +4
Safety cases for frontier AI systems should provide a convincing argument, supported by evidence, that the risk of harm is within an acceptable bound. When developers author their…