LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
arXiv:2504.19076 · doi:10.1145/3731120.3744588
Abstract
Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based evaluations often align with human judgments, leading some to suggest that human judges may no longer be necessary, while others highlight concerns about judgment reliability, validity, and long-term impact. As IR systems begin incorporating LLM-generated signals, evaluation outcomes risk becoming self-reinforcing, potentially leading to misleading conclusions. This paper examines scenarios where LLM-evaluators may falsely indicate success, particularly when LLM-based judgments influence both system development and evaluation. We highlight key risks, including bias reinforcement, reproducibility challenges, and inconsistencies in assessment methodologies. To address these concerns, we propose tests to quantify adverse effects, guardrails, and a collaborative framework for constructing reusable test collections that integrate LLM judgments responsibly. By providing perspectives from academia and industry, this work aims to establish best practices for the principled use of LLMs in IR evaluation.
References in corpus (10)
- Perspectives on Large Language Models for Relevance Judgment
- The Information Retrieval Experiment Platform
- Neural Retrievers are Biased Towards LLM-Generated Content
- On the Evaluation of Machine-Generated Reports
- Do LLMs Memorize Recommendation Datasets? A Preliminary Study on MovieLens-1M
- LLM Evaluators Recognize and Favor Their Own Generations
- Extrinsic Evaluation of Cultural Competence in Large Language Models
- Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I
- Evaluating Large Language Model Biases in Persona-Steered Generation
- Self-training Large Language Models through Knowledge Detection