3 papers
cs.HC2026
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Zeyu He, Xuan Qi, Subramanian Chidambaram +4
Large language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a n…
cs.MA2026
How to Steer Your Multi-Agent System: Human-LLM Collaborative Planning
Zeyu He, Hannah Kim, Dan Zhang +1
In orchestrated multi-agent systems, humans often struggle to manage plans due to their complexity and limited transparency. Existing approaches rely on outcome-level supervision,…
cs.HC2025
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
Zeyu He, Saniya Naphade, Ting-Hao 'Kenneth' Huang
Millions of users prompt large language models (LLMs) for various tasks, but how good are people at prompt engineering? Do users actually get closer to their desired outcome over m…