4 papers
Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design
Leon Eshuijs, Shihan Wang, Antske Fokkens
Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the conditions under which this occur…
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
Leon Eshuijs, Archie Chaudhury, Alan McBeth +1
LLM-as-a-judge is widely used as a scalable substitute for human evaluation, yet current approaches rely on black-box access and struggle to detect subtle dishonesty, such as sycop…
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
Leon Eshuijs, Shihan Wang, Antske Fokkens
Reliance on spurious correlations (shortcuts) has been shown to underlie many of the successes of language models. Previous work focused on identifying the input elements that impa…
Automatic Evaluation Metrics for Artificially Generated Scientific Research
Niklas Höpner, Leon Eshuijs, Dimitrios Alivanistos +2
Foundation models are increasingly used in scientific research, but evaluating AI-generated scientific work remains challenging. While expert reviews are costly, large language mod…