6 citations · 6 across the 5 of their papers we have counts for
5 papers · 1 filter
Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
Oliver Daniels, Perusha Moodley, Benjamin M. Marlin +1
Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is…
Exploration Hacking: Can LLMs Learn to Resist RL Training?
Eyon Jang, Damon Falck, Joschka Braun +6
Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignment. Successful RL relies on suf…
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
Oliver Daniels, Perusha Moodley, Benjamin M. Marlin +1
Alignment audits aim to robustly identify hidden goals from strategic, situationally aware misaligned models. Despite this threat model, existing auditing methods have not been sys…
Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Action Spaces
Perusha Moodley, Pramod Kaushik, Dhillu Thambi +4
Decision Transformers, in their vanilla form, struggle to perform on image-based environments with multi-discrete action spaces. Although enhanced Decision Transformer architecture…
A Conservative Q-Learning approach for handling distribution shift in sepsis treatment strategies
Pramod Kaushik, Sneha Kummetha, Perusha Moodley +1
Sepsis is a leading cause of mortality and its treatment is very expensive. Sepsis treatment is also very challenging because there is no consensus on what interventions work best…