3 papers
cs.CY2026
Detecting Offensive Cyber Agents: A Detection-in-Depth Approach
Matt Mittelsteadt, Jam Kraprayoon, Robin Staes-Polet +4
Artificial Intelligence (AI) agents can now orchestrate cyberattacks. This development is already increasing the speed and scale of cyber attacks, decreasing attack costs, and impr…
cs.CL2024
Representation Noising: A Defence Mechanism Against Harmful Finetuning
Domenic Rosati, Jan Wehner, Kai Williams +7
Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release o…
cs.CL2024
Immunization against harmful fine-tuning attacks
Domenic Rosati, Jan Wehner, Kai Williams +4
Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM o…