6 papers
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
Hyo Jin Do, Zahra Ashktorab, Jasmina Gajcin +5
The LLM-as-a-judge paradigm enables flexible, user-defined evaluation, but its effectiveness is often limited by the scarcity of diverse, representative data for refining criteria.…
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
Zahra Ashktorab, Werner Geyer, Michael Desmond +6
With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given tas…
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
Zahra Ashktorab, Michael Desmond, Qian Pan +7
Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes ti…
Granite Guardian
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +20
We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with…
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
Nico Wagner, Michael Desmond, Rahul Nair +6
LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty…
Human-Centered Design Recommendations for LLM-as-a-Judge
Qian Pan, Zahra Ashktorab, Michael Desmond +5
Traditional reference-based metrics, such as BLEU and ROUGE, are less effective for assessing outputs from Large Language Models (LLMs) that produce highly creative or superior-qua…