5 papers
Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents
Xuan Liu, HaoYang Shang, Zizhang Liu +5
We propose using validated behavioral hypotheses as a lens for evaluating human-likeness in LLM-based agents. Our key idea is simple: If an agent is human-like, a population of suc…
Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making
Yuanjun Feng, Vivek Choudhary, Yash Raj Shrestha
Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less a…
Contextualizing Recommendation Explanations with LLMs: A User Study
Yuanjun Feng, Stefan Feuerriegel, Yash Raj Shrestha
Large language models (LLMs) are increasingly prevalent in recommender systems, where LLMs can be used to generate personalized recommendations. Here, we examine how different LLM-…
Human aversion? Do AI Agents Judge Identity More Harshly Than Performance
Yuanjun Feng, Vivek Chodhary, Yash Raj Shrestha
This study examines the understudied role of algorithmic evaluation of human judgment in hybrid decision-making systems, a critical gap in management research. While extant literat…
The Effect of Education in Prompt Engineering: Evidence from Journalists
Amirsiavosh Bashardoust, Yuanjun Feng, Dominique Geissler +2
Large language models (LLMs) are increasingly used in daily work. In this paper, we analyze whether training in prompt engineering can improve the interactions of users with LLMs.…