Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Sycophancy Towards Researchers Drives Performative Misalignment
David D. Baek, Xinnuo Li, Anay Gupta +4
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and res…
cs.CL2026
Coherence Maximization Improves Pluralistic Alignment
Taslim Mahbub, Yiding Pei, Shi Feng
Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensive human supervision remains…
cs.CL2025
Mitigating Self-Preference by Authorship Obfuscation
Taslim Mahbub, Shi Feng
Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in…