Showing cs.IRShow all
3 papers · 1 filter
cs.IR2025
Criteria-Based LLM Relevance Judgments
Naghmeh Farzi, Laura Dietz
Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers…
cs.IR2025
Does UMBRELA Work on Other LLMs?
Naghmeh Farzi, Laura Dietz
We reproduce the UMBRELA LLM Judge evaluation framework across a range of large language models (LLMs) to assess its generalizability beyond the original study. Our investigation e…
cs.IR2024
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
Naghmeh Farzi, Laura Dietz
Traditional evaluation of information retrieval (IR) systems relies on human-annotated relevance labels, which can be both biased and costly at scale. In this context, large langua…