13 papers
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
Charles Chiang, Simret Gebreegziabher, Annalisa Szymanski +6
LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding t…
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
Jaeyoon Song, Zahra Ashktorab, Thomas W. Malone
Scheduling is a perennial-and often challenging-problem for many groups. Existing tools are mostly static, showing an identical set of choices to everyone, regardless of the curren…
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
Simret Araya Gebreegziabher, Yukun Yang, Charles Chiang +7
Large Language Model (LLM)-powered web GUI agents are increasingly automating everyday online tasks. Despite their popularity, little is known about how users' preferences and valu…
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
Hyo Jin Do, Zahra Ashktorab, Jasmina Gajcin +5
The LLM-as-a-judge paradigm enables flexible, user-defined evaluation, but its effectiveness is often limited by the scarcity of diverse, representative data for refining criteria.…
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
Zahra Ashktorab, Werner Geyer, Michael Desmond +6
With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given tas…
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
Zahra Ashktorab, Michael Desmond, Qian Pan +7
Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes ti…