3 papers
cs.CL2026
Supporting Humans in Evaluating AI Summaries of Legal Depositions
Naghmeh Farzi, Laura Dietz, Dave D. Lewis
While large language models (LLMs) are increasingly used to summarize long documents, this trend poses significant challenges in the legal domain, where the factual accuracy of dep…
cs.IR2025
Criteria-Based LLM Relevance Judgments
Naghmeh Farzi, Laura Dietz
Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers…
cs.IR2025
Does UMBRELA Work on Other LLMs?
Naghmeh Farzi, Laura Dietz
We reproduce the UMBRELA LLM Judge evaluation framework across a range of large language models (LLMs) to assess its generalizability beyond the original study. Our investigation e…