1 paper · 1 filter
Binghai Wang, Yantao Liu, Yuxuan Liu +13
Generative Reward Models (GenRMs) and LLM-as-a-Judge exhibit deceptive alignment by producing correct judgments for incorrect reasons, as they are trained and evaluated to prioriti…