activity
20242026
collaborators

6 papers

cs.IR2026

Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations

Laura Dietz

Evaluating AI/Agentic system outputs reliably requires human judgment, but how one incorporates the human determines whether one gets a real quality signal or expensive theater. Th…

cs.CL2026

Supporting Humans in Evaluating AI Summaries of Legal Depositions

Naghmeh Farzi, Laura Dietz, Dave D. Lewis

While large language models (LLMs) are increasingly used to summarize long documents, this trend poses significant challenges in the legal domain, where the factual accuracy of dep…

cs.IR2026

LLM-based relevance assessment still can't replace human relevance assessment

Charles L. A. Clarke, Laura Dietz

The use of large language models (LLMs) for relevance assessment in information retrieval has gained significant attention, with recent studies suggesting that LLM-based judgments…

cs.IR2025

Criteria-Based LLM Relevance Judgments

Naghmeh Farzi, Laura Dietz

Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers…

cs.IR2025

Does UMBRELA Work on Other LLMs?

Naghmeh Farzi, Laura Dietz

We reproduce the UMBRELA LLM Judge evaluation framework across a range of large language models (LLMs) to assess its generalizability beyond the original study. Our investigation e…

cs.IR2024

Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3

Naghmeh Farzi, Laura Dietz

Traditional evaluation of information retrieval (IR) systems relies on human-annotated relevance labels, which can be both biased and costly at scale. In this context, large langua…