5 papers · 1 filter
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Zeyu He, Xuan Qi, Subramanian Chidambaram +4
Large language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a n…
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
Zeyu He, Saniya Naphade, Ting-Hao 'Kenneth' Huang
Millions of users prompt large language models (LLMs) for various tasks, but how good are people at prompt engineering? Do users actually get closer to their desired outcome over m…
What Color Scheme is More Effective in Assisting Readers to Locate Information in a Color-Coded Article?
Ho Yin Ng, Zeyu He, Ting-Hao 'Kenneth' Huang
Color coding, a technique assigning specific colors to cluster information types, has proven advantages in aiding human cognitive activities, especially reading and comprehension.…
If in a Crowdsourced Data Annotation Pipeline, a GPT-4
Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding +2
Recent studies indicated GPT-4 outperforms online crowd workers in data labeling accuracy, notably workers from Amazon Mechanical Turk (MTurk). However, these studies were criticiz…
How Does Conversation Length Impact User's Satisfaction? A Case Study of Length-Controlled Conversations with LLM-Powered Chatbots
Shih-Hong Huang, Ya-Fang Lin, Zeyu He +2
Users can discuss a wide range of topics with large language models (LLMs), but they do not always prefer solving problems or getting information through lengthy conversations. Thi…