6 papers
Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations
Laura Dietz
Evaluating AI/Agentic system outputs reliably requires human judgment, but how one incorporates the human determines whether one gets a real quality signal or expensive theater. Th…
Supporting Humans in Evaluating AI Summaries of Legal Depositions
Naghmeh Farzi, Laura Dietz, Dave D. Lewis
While large language models (LLMs) are increasingly used to summarize long documents, this trend poses significant challenges in the legal domain, where the factual accuracy of dep…
LLM-based relevance assessment still can't replace human relevance assessment
Charles L. A. Clarke, Laura Dietz
The use of large language models (LLMs) for relevance assessment in information retrieval has gained significant attention, with recent studies suggesting that LLM-based judgments…
Criteria-Based LLM Relevance Judgments
Naghmeh Farzi, Laura Dietz
Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers…
Does UMBRELA Work on Other LLMs?
Naghmeh Farzi, Laura Dietz
We reproduce the UMBRELA LLM Judge evaluation framework across a range of large language models (LLMs) to assess its generalizability beyond the original study. Our investigation e…
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
Naghmeh Farzi, Laura Dietz
Traditional evaluation of information retrieval (IR) systems relies on human-annotated relevance labels, which can be both biased and costly at scale. In this context, large langua…