1 paper · 1 filter
Tim Tian Hua, Andrew Qin, Samuel Marks +1
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations…