2 papers
cs.CY2026
Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents
Xuan Liu, HaoYang Shang, Zizhang Liu +5
We propose using validated behavioral hypotheses as a lens for evaluating human-likeness in LLM-based agents. Our key idea is simple: If an agent is human-like, a population of suc…
cs.AI2026
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
Xuan Liu, Haoyang Shang, Zizhang Liu +4
Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choi…