3 papers
cs.AI2026
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
Akshat Naik, Emma Gouné, Patrick Quinn +4
As Large Language Model (LLM) agents become more widespread, associated misalignment risks increase. While prior research has studied agents' ability to produce harmful outputs or…
cs.CL2025
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
Jason R Brown, Lennie Wells, Edward James Young +1
Proximal Policy Optimisation (PPO) is an established and effective policy gradient algorithm used for Language Model Reinforcement Learning from Human Feedback (LM-RLHF). PPO perfo…
cs.AI2025
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
Rishane Dassanayake, Mario Demetroudi, James Walpole +3
Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasio…