1 paper · 1 filter
Zhengyu Hu, Linxin Song, Jieyu Zhang +7
The use of large language models (LLMs) as judges, particularly in preference comparisons, has become widespread, but this reveals a notable bias towards longer responses, undermin…