9 papers
Stayin' Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data
Simret Araya Gebreegziabher, Allison E Sproul, Yinuo Yang +3
Current human-AI alignment and evaluation methods for large language models (LLMs) often rely on preference signals collected immediately after an interaction. This practice implic…
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
Charles Chiang, Simret Gebreegziabher, Annalisa Szymanski +6
LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding t…
Comparing Human Oversight Strategies for Computer-Use Agents
Chaoran Chen, Zhiping Zhang, Zeya Chen +9
LLM-powered computer-use agents (CUAs) are shifting users from direct manipulation to supervisory coordination. Existing oversight mechanisms, however, have largely been studied as…
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
Simret Araya Gebreegziabher, Yukun Yang, Charles Chiang +7
Large Language Model (LLM)-powered web GUI agents are increasingly automating everyday online tasks. Despite their popularity, little is known about how users' preferences and valu…
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
Jingyu Tang, Chaoran Chen, Jiawen Li +11
The dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the ris…
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov +7
The rise of Large Language Models (LLMs) has revolutionized Graphical User Interface (GUI) automation through LLM-powered GUI agents, yet their ability to process sensitive data wi…