activity
20242026
collaborators

8 papers

cs.HC2026

MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria

Charles Chiang, Simret Gebreegziabher, Annalisa Szymanski +6

LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding t…

cs.HC2026

"Better Ask for Forgiveness than Permission": Practices and Policies of AI Disclosure in Freelance Work

Angel Hsing-Chi Hwang, Senya Wong, Baixiao Chen +2

The growing use of AI applications among freelance workers is reshaping trust and relationships with clients. This paper investigates how both workers and clients perceive AI use a…

cs.HC2026

The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes

Simret Araya Gebreegziabher, Yukun Yang, Charles Chiang +7

Large Language Model (LLM)-powered web GUI agents are increasingly automating everyday online tasks. Despite their popularity, little is known about how users' preferences and valu…

cs.HC2025

Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges

Hyo Jin Do, Zahra Ashktorab, Jasmina Gajcin +5

The LLM-as-a-judge paradigm enables flexible, user-defined evaluation, but its effectiveness is often limited by the scarcity of diverse, representative data for refining criteria.…

cs.HC2025

Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust

Hyo Jin Do, Werner Geyer

Large language models are known to produce outputs that are plausible but factually incorrect. To prevent people from making erroneous decisions by blindly trusting AI, researchers…

cs.HC2025

Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators

Hyo Jin Do, Rachel Ostrand, Werner Geyer +3

Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advan…