activity
20242026
most citedLLM Evaluators Recognize and Favor Their Own Generations

15 citations · 31 across the 15 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2026

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

Jesse St. Amand, Callum Canavan, Sohaib Imran +5

Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors: an LLM may reco…

cs.CL2026

What are Key Factors for Updates in RL for LLM Reasoning?

Peidong Wang, Demi Wang, Xufang Luo +5

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existi…

cs.CL2026

Building Comparative Motivation Profiles with Instrumental Interventions

David Vella Zarb, Rustem Turtayev, Taywon Min +2

Safety evaluations often infer latent motivations from behavioral patterns, but the construct validity of these inferences is unclear. We study this problem in alignment faking, wh…

cs.CL2026

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

Arush Tagade, Shaoheng Zhou, Jiaxin Wen +1

Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's…

cs.CL2025

Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment

Manas Khatore, Sumana Sridharan, Kevork Sulahian +2

Automated answer matching, which leverages LLMs to evaluate free-text responses by comparing them to a reference answer, shows substantial promise as a scalable and aligned alterna…

cs.CL2025

Mitigating Self-Preference by Authorship Obfuscation

Taslim Mahbub, Shi Feng

Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in…