12 papers
CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories
Zheyuan Deng, Binghang Lu, Hanqi Feng +10
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet…
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Sky Ng, Brihi Joshi, Ishan Gupta +47
Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain i…
PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications
Yifan Simon Liu, Qianfeng Wen, Yilan Fan +40
Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Hanwen Xing, Pengyun Wang, BingXu Meng +8
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
Yuexing Hao, Xiaomin Li
Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downs…