8 citations · 16 across the 4 of their papers we have counts for
4 papers · 1 filter
Criteria-Based LLM Relevance Judgments
Naghmeh Farzi, Laura Dietz
Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers…
Does UMBRELA Work on Other LLMs?
Naghmeh Farzi, Laura Dietz
We reproduce the UMBRELA LLM Judge evaluation framework across a range of large language models (LLMs) to assess its generalizability beyond the original study. Our investigation e…
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
Naghmeh Farzi, Laura Dietz
Traditional evaluation of information retrieval (IR) systems relies on human-annotated relevance labels, which can be both biased and costly at scale. In this context, large langua…
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
Naghmeh Farzi, Laura Dietz
Current IR evaluation is based on relevance judgments, created either manually or automatically, with decisions outsourced to Large Language Models (LLMs). We offer an alternative…