5 papers · 1 filter
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these enviro…
Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design
Leon Eshuijs, Shihan Wang, Antske Fokkens
Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the conditions under which this occur…
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
Leon Eshuijs, Shihan Wang, Antske Fokkens
Reliance on spurious correlations (shortcuts) has been shown to underlie many of the successes of language models. Previous work focused on identifying the input elements that impa…
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
Leon Eshuijs, Archie Chaudhury, Alan McBeth +1
LLM-as-a-judge is widely used as a scalable substitute for human evaluation, yet current approaches rely on black-box access and struggle to detect subtle dishonesty, such as sycop…
Balancing the Scales: Reinforcement Learning for Fair Classification
Leon Eshuijs, Shihan Wang, Antske Fokkens
Fairness in classification tasks has traditionally focused on bias removal from neural representations, but recent trends favor algorithmic methods that embed fairness into the tra…