From the 1 of 12 linked papers with an AI index.
1 paper · 1 filter
Xing Zhang, Guanghui Wang, Yanwei Cui +4
Telling an LLM to "be enthusiastic" raises its sycophancy rate from 30\% to 50\% on a lightly-aligned model, but has zero effect on a strongly-aligned one. We define this gap as th…