5 papers
Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation
Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie +1
Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment re…
LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal
Pietro Bernardelle, Samaneh Mohtadi, Stefano Civelli +2
Large language models (LLMs) are increasingly used in information retrieval (IR) pipelines as relevance judges and re-rankers. Yet most analyses remain output-centric, evaluating g…
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
Chuting Yu, Hang Li, Guido Zuccon +2
Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using larg…
Revisiting Human-vs-LLM judgments using the TREC Podcast Track
Watheq Mansour, J. Shane Culpepper, Joel Mackenzie +1
Using large language models (LLMs) to annotate relevance is an increasingly important technique in the information retrieval community. While some studies demonstrate that LLMs can…
Reassessing Collaborative Writing Theories and Frameworks in the Age of LLMs: What Still Applies and What We Must Leave Behind
Daisuke Yukita, Tim Miller, Joel Mackenzie
In this paper, we conduct a critical review of existing theories and frameworks on human-human collaborative writing to assess their relevance to the current human-AI paradigm in o…