collaborators

5 papers

cs.HC2025

Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges

Hyo Jin Do, Zahra Ashktorab, Jasmina Gajcin +5

The LLM-as-a-judge paradigm enables flexible, user-defined evaluation, but its effectiveness is often limited by the scarcity of diverse, representative data for refining criteria.…

cs.HC2025

EvalAssist: A Human-Centered Tool for LLM-as-a-Judge

Zahra Ashktorab, Werner Geyer, Michael Desmond +6

With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given tas…

cs.CL2024

Granite Guardian

Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +20

We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with…

cs.LG2024

Black-box Uncertainty Quantification Method for LLM-as-a-Judge

Nico Wagner, Michael Desmond, Rahul Nair +6

LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty…

cs.HC2024

Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences

Zahra Ashktorab, Michael Desmond, Qian Pan +7

Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes ti…