15 citations · 31 across the 15 of their papers we have counts for
12 papers · 1 filter
Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
Jesse St. Amand, Callum Canavan, Sohaib Imran +5
Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors: an LLM may reco…
What are Key Factors for Updates in RL for LLM Reasoning?
Peidong Wang, Demi Wang, Xufang Luo +5
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existi…
Building Comparative Motivation Profiles with Instrumental Interventions
David Vella Zarb, Rustem Turtayev, Taywon Min +2
Safety evaluations often infer latent motivations from behavioral patterns, but the construct validity of these inferences is unclear. We study this problem in alignment faking, wh…
Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment
Arush Tagade, Shaoheng Zhou, Jiaxin Wen +1
Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's…
Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment
Manas Khatore, Sumana Sridharan, Kevork Sulahian +2
Automated answer matching, which leverages LLMs to evaluate free-text responses by comparing them to a reference answer, shows substantial promise as a scalable and aligned alterna…
Mitigating Self-Preference by Authorship Obfuscation
Taslim Mahbub, Shi Feng
Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in…