1 paper
Pranav Mahajan, Ihor Kendiukhov, Syed Hussain +1
Recent work identifies a stated-revealed (SvR) preference gap in language models (LMs): a mismatch between the values models endorse and the choices they make in context. Existing…