3 papers
cs.IR2026
The Effect of Multi-Lingual and Keyword Adversarial Injection on LLM Relevance Judgment
Nguyen Khoi Vo, Duy Duong Tuong, Oleg Zendel +1
Large language models (LLMs) are increasingly being used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation rem…
cs.IR2025
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
Laura Dietz, Oleg Zendel, Peter Bailey +6
Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empi…
cs.IR2025
Generative Information Retrieval Evaluation
Marwah Alaofi, Negar Arabzadeh, Charles L. A. Clarke +1
In this chapter, we consider generative information retrieval evaluation from two distinct but interrelated perspectives. First, large language models (LLMs) themselves are rapidly…