collaborators

6 papers

cs.LG2026

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment

Jason R. Brown, Patrick Leask, Lev McKinney

Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned beh…

cs.AI2026

Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff +18

Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning intern…

cs.AI2026

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

Akshat Naik, Emma Gouné, Patrick Quinn +4

As Large Language Model (LLM) agents become more widespread, associated misalignment risks increase. While prior research has studied agents' ability to produce harmful outputs or…

cs.LG2026

Understanding Goal Generalisation in Sequential Reinforcement Learning

Jason Ross Brown, Edward James Young

Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled understanding of how such age…

cs.CL2025

KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF

Jason R Brown, Lennie Wells, Edward James Young +1

Proximal Policy Optimisation (PPO) is an established and effective policy gradient algorithm used for Language Model Reinforcement Learning from Human Feedback (LM-RLHF). PPO perfo…

cs.AI2025

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

Rishane Dassanayake, Mario Demetroudi, James Walpole +3

Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasio…