3 papers
cs.IR2026
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
Chuting Yu, Hang Li, Guido Zuccon +2
Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using larg…
cs.IR2026
Revisiting Human-vs-LLM judgments using the TREC Podcast Track
Watheq Mansour, J. Shane Culpepper, Joel Mackenzie +1
Using large language models (LLMs) to annotate relevance is an increasingly important technique in the information retrieval community. While some studies demonstrate that LLMs can…
cs.HC2025
Reassessing Collaborative Writing Theories and Frameworks in the Age of LLMs: What Still Applies and What We Must Leave Behind
Daisuke Yukita, Tim Miller, Joel Mackenzie
In this paper, we conduct a critical review of existing theories and frameworks on human-human collaborative writing to assess their relevance to the current human-AI paradigm in o…