collaborators
Showing cs.HCShow all

5 papers · 1 filter

cs.HC2025

Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges

Hyo Jin Do, Zahra Ashktorab, Jasmina Gajcin +5

The LLM-as-a-judge paradigm enables flexible, user-defined evaluation, but its effectiveness is often limited by the scarcity of diverse, representative data for refining criteria.…

cs.HC2025

EvalAssist: A Human-Centered Tool for LLM-as-a-Judge

Zahra Ashktorab, Werner Geyer, Michael Desmond +6

With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given tas…

cs.HC2025

Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences

Zahra Ashktorab, Michael Desmond, Qian Pan +7

Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes ti…

cs.HC2025

Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations

Zahra Ashktorab, Qian Pan, Werner Geyer +5

We investigate the impact of hallucinations and Cognitive Forcing Functions in human-AI collaborative content-grounded data generation, focusing on the use of Large Language Models…

cs.HC2024

Human-Centered Design Recommendations for LLM-as-a-Judge

Qian Pan, Zahra Ashktorab, Michael Desmond +5

Traditional reference-based metrics, such as BLEU and ROUGE, are less effective for assessing outputs from Large Language Models (LLMs) that produce highly creative or superior-qua…