2 papers
cs.CR2025
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
Kai Williams, Rohan Subramani, Francis Rhys Ward
Frontier AI developers may fail to align or control highly-capable AI agents. In many cases, it could be useful to have emergency shutdown mechanisms which effectively prevent misa…
cs.CL2024
Representation Noising: A Defence Mechanism Against Harmful Finetuning
Domenic Rosati, Jan Wehner, Kai Williams +7
Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release o…