12 papers
Exploring the Rashomon Set for Concept-Based Models
Shihan Feng, Cheng Zhang, Michael Xi +3
In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fundamentally different internal logic. How…
What are Key Factors for Updates in RL for LLM Reasoning?
Peidong Wang, Demi Wang, Xufang Luo +5
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existi…
Building Comparative Motivation Profiles with Instrumental Interventions
David Vella Zarb, Rustem Turtayev, Taywon Min +2
Safety evaluations often infer latent motivations from behavioral patterns, but the construct validity of these inferences is unclear. We study this problem in alignment faking, wh…
Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment
Arush Tagade, Shaoheng Zhou, Jiaxin Wen +1
Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's…
Self-Improvement as Coherence Optimization: A Theoretical Account
Tianyi Qiu, Ahmed Hani Ismail, Zhonghao He +1
Can language models improve their accuracy without external supervision? Methods such as debate, bootstrap, and internal coherence maximization achieve this surprising feat, even m…
Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment
Manas Khatore, Sumana Sridharan, Kevork Sulahian +2
Automated answer matching, which leverages LLMs to evaluate free-text responses by comparing them to a reference answer, shows substantial promise as a scalable and aligned alterna…