10 papers
Agentopia: Long-Term Life Simulation and Learning in Agent Societies
Xintao Wang, Sirui Zheng, Hongqiu Wu +10
Humans learn from social life. Simulating this process with LLM-powered agents represents a promising research direction, raising a natural question: whether LLMs can learn from su…
You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
Nicole Hsing, Asuka Yuxi Zheng, Yi Zhao +2
Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. We show that a single aligned a…
Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making
Jen-tse Huang, Didi Zhou, Faith Kamau +5
Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as clinical decision support and medical documentation. However, the robustness of these models a…
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
Man Ho Lam, Chaozheng Wang, Hange Liu +5
Coding agents powered by large language models are increasingly expected to perform realistic software maintenance tasks beyond isolated issue resolution. Existing benchmarks have…
How to Interpret Agent Behavior
Jie Gao, Kaiser Sun, Jen-tse Huang +8
Autonomous agents such as Claude Code and Codex now operate for hours or even days. Understanding their runtime behavior has become critical for downstream tasks such as diagnosing…
The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
Yunze Xiao, Vivienne J. Zhang, Chenghao Yang +3
Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{P…