2 papers
cs.CL2026
Sycophancy Towards Researchers Drives Performative Misalignment
David D. Baek, Xinnuo Li, Anay Gupta +4
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and res…
cs.CL2024
ReIFE: Re-evaluating Instruction-Following Evaluation
Yixin Liu, Kejian Shi, Alexander R. Fabbri +5
The automatic evaluation of instruction following typically involves using large language models (LLMs) to assess response quality. However, there is a lack of comprehensive evalua…