2 citations · 3 across the 4 of their papers we have counts for
3 papers · 1 filter
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
Hyo Jin Do, Zahra Ashktorab, Jasmina Gajcin +5
The LLM-as-a-judge paradigm enables flexible, user-defined evaluation, but its effectiveness is often limited by the scarcity of diverse, representative data for refining criteria.…
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
Zahra Ashktorab, Werner Geyer, Michael Desmond +6
With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given tas…
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
Zahra Ashktorab, Michael Desmond, Qian Pan +7
Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes ti…