Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
Ziyi Wang, Yuxuan Lu, Wenbo Li +13
Can large language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generating ``believable'' human behavio…
cs.CL2026
NARRA-Gym for Evaluating Interactive Narrative Agents
Yue Huang, Yuchen Ma, Jiayi Ye +14
Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limit…
cs.CL2025
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
Yuxuan Lu, Bingsheng Yao, Hansu Gu +7
Usability testing is a fundamental research method that user experience (UX) researchers use to evaluate and iterate their new designs. But what about evaluating and iterating the…