4 papers
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
Charles Chiang, Simret Gebreegziabher, Annalisa Szymanski +6
LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding t…
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
Simret Araya Gebreegziabher, Yukun Yang, Charles Chiang +7
Large Language Model (LLM)-powered web GUI agents are increasingly automating everyday online tasks. Despite their popularity, little is known about how users' preferences and valu…
Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
Hyo Jin Do, Werner Geyer
Large language models are known to produce outputs that are plausible but factually incorrect. To prevent people from making erroneous decisions by blindly trusting AI, researchers…
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
Hyo Jin Do, Rachel Ostrand, Werner Geyer +3
Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advan…