collaborators

12 papers

cs.LG2026

Exploring the Rashomon Set for Concept-Based Models

Shihan Feng, Cheng Zhang, Michael Xi +3

In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fundamentally different internal logic. How…

cs.CL2026

What are Key Factors for Updates in RL for LLM Reasoning?

Peidong Wang, Demi Wang, Xufang Luo +5

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existi…

cs.CL2026

Building Comparative Motivation Profiles with Instrumental Interventions

David Vella Zarb, Rustem Turtayev, Taywon Min +2

Safety evaluations often infer latent motivations from behavioral patterns, but the construct validity of these inferences is unclear. We study this problem in alignment faking, wh…

cs.CL2026

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

Arush Tagade, Shaoheng Zhou, Jiaxin Wen +1

Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's…

cs.LG2026

Self-Improvement as Coherence Optimization: A Theoretical Account

Tianyi Qiu, Ahmed Hani Ismail, Zhonghao He +1

Can language models improve their accuracy without external supervision? Methods such as debate, bootstrap, and internal coherence maximization achieve this surprising feat, even m…

cs.CL2025

Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment

Manas Khatore, Sumana Sridharan, Kevork Sulahian +2

Automated answer matching, which leverages LLMs to evaluate free-text responses by comparing them to a reference answer, shows substantial promise as a scalable and aligned alterna…