4 papers
Sycophancy Towards Researchers Drives Performative Misalignment
David D. Baek, Xinnuo Li, Anay Gupta +4
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and res…
Coherence Maximization Improves Pluralistic Alignment
Taslim Mahbub, Yiding Pei, Shi Feng
Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensive human supervision remains…
Mitigating Self-Preference by Authorship Obfuscation
Taslim Mahbub, Shi Feng
Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in…
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
Khizar Anjum, Muhammad Arbab Arshad, Kadhim Hayawi +10
Large language models (LLMs) are increasingly being deployed across disciplines due to their advanced reasoning and problem solving capabilities. To measure their effectiveness, va…