3 papers
cs.AI2026
AgentPersonaBench: Benchmarking Persona-Driven User Simulation
Jintao Huang, Yifan Wang, Hongyu Shen +43
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deploy…
cs.CL2026
Can Language Models Learn to Forecast Stock Prices
Jiacheng Guo, Suozhi Huang, Shuzhen Li +11
Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and co…
cs.AI2026
Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
Weiyuan Li, Jinghan Xu, Aili Chen +4
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interac…