2 papers
cs.LG2025
Likert or Not: LLM Absolute Relevance Judgments on Fine-Grained Ordinal Scales
Charles Godfrey, Ping Nie, Natalia Ostapuk +3
Large language models (LLMs) obtain state of the art zero shot relevance ranking performance on a variety of information retrieval tasks. The two most common prompts to elicit LLM…
cs.CL2024
Measuring the Groundedness of Legal Question-Answering Systems
Dietrich Trautmann, Natalia Ostapuk, Quentin Grail +4
In high-stakes domains like legal question-answering, the accuracy and trustworthiness of generative AI systems are of paramount importance. This work presents a comprehensive benc…