1 paper · 1 filter
Riya Tapwal, Abhishek Kumar, Carsten Maple
Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such as position, verbosity, and…