Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgentPersonaBench: Benchmarking Persona-Driven User Simulation
Jintao Huang, Yifan Wang, Hongyu Shen +43
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deploy…
cs.AI2026
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…