collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs

Kyle O'Brien, Stephen Casper, Quentin Anthony +7

Open-weight AI systems offer unique benefits, including enhanced transparency, open research, and decentralized access. However, they are vulnerable to tampering attacks which can…

cs.LG2025

Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs

Xander Davies, Eric Winsor, Alexandra Souly +4

LLM developers have imposed technical interventions to prevent fine-tuning misuse attacks, attacks where adversaries evade safeguards by fine-tuning the model using a public API. P…

cs.LG2025

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Alexandra Souly, Javier Rando, Ed Chapman +10

Poisoning attacks can compromise the safety of large language models (LLMs) by injecting malicious documents into their training data. Existing work has studied pretraining poisoni…

cs.LG2025

Reward Model Overoptimisation in Iterated RLHF

Lorenz Wolf, Robert Kirk, Mirco Musolesi

Reinforcement learning from human feedback (RLHF) is a widely used method for aligning large language models with human preferences. However, RLHF often suffers from reward model o…

cs.LG2025

Existing Large Language Model Unlearning Evaluations Are Inconclusive

Zhili Feng, Yixuan Even Xu, Alexander Robey +5

Machine unlearning aims to remove sensitive or undesired data from large language models. However, recent studies suggest that unlearning is often shallow, claiming that removed kn…

cs.LG2025

An Example Safety Case for Safeguards Against Misuse

Joshua Clymer, Jonah Weinbaum, Robert Kirk +3

Existing evaluations of AI misuse safeguards provide a patchwork of evidence that is often difficult to connect to real-world decisions. To bridge this gap, we describe an end-to-e…