2 papers
cs.AI2026
AgentPersonaBench: Benchmarking Persona-Driven User Simulation
Jintao Huang, Yifan Wang, Hongyu Shen +43
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deploy…
cs.AI2026
COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement
Siwei Chen, Xinping Bao, Xinyu Cai +3
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Ex…