activity
20242026
collaborators

5 papers

cs.CY2026

Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents

Xuan Liu, HaoYang Shang, Zizhang Liu +5

We propose using validated behavioral hypotheses as a lens for evaluating human-likeness in LLM-based agents. Our key idea is simple: If an agent is human-like, a population of suc…

cs.CE2025

Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making

Yuanjun Feng, Vivek Choudhary, Yash Raj Shrestha

Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less a…

cs.HC2025

Contextualizing Recommendation Explanations with LLMs: A User Study

Yuanjun Feng, Stefan Feuerriegel, Yash Raj Shrestha

Large language models (LLMs) are increasingly prevalent in recommender systems, where LLMs can be used to generate personalized recommendations. Here, we examine how different LLM-…

cs.HC2025

Human aversion? Do AI Agents Judge Identity More Harshly Than Performance

Yuanjun Feng, Vivek Chodhary, Yash Raj Shrestha

This study examines the understudied role of algorithmic evaluation of human judgment in hybrid decision-making systems, a critical gap in management research. While extant literat…

cs.HC2024

The Effect of Education in Prompt Engineering: Evidence from Journalists

Amirsiavosh Bashardoust, Yuanjun Feng, Dominique Geissler +2

Large language models (LLMs) are increasingly used in daily work. In this paper, we analyze whether training in prompt engineering can improve the interactions of users with LLMs.…