1 paper · 1 filter
Yang Zhang, Cunxiang Wang, Lindong Wu +4
Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own.…