From the 2 of 7 linked papers with an AI index.
7 papers
Multi-Agent LLMs Fail to Explore Each Other
Hyeong Kyu Choi, Jiatong Li, Wendi Li +2
The paper shows that large language model agents struggle to explore each other in multi-agent settings, leading to poor coordination, and introduces the MACE framework that uses s…
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33
The paper presents OSWorld 2.0, a benchmark consisting of 108 long‑horizon, real‑world computer‑use workflows designed to evaluate how well AI agents can handle complex, multi‑step…
On the Reliability of Computer Use Agents
Gonzalo Gonzalez-Pumariega, Saaket Agashe, Jiachen Yang +2
Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet…
Scaling Agents for Computer Use
Gonzalo Gonzalez-Pumariega, Vincent Tu, Chih-Lun Lee +3
Computer-use agents (CUAs) hold promise for automating everyday digital tasks, but their performance on long-horizon, complex problems remains unreliable. Single-rollout execution…
Agents of Change: Self-Evolving LLM Agents for Strategic Planning
Nikolas Belle, Dakota Barnes, Alfonso Amayuelas +3
We address the long-horizon gap in large language model (LLM) agents by enabling them to sustain coherent strategies in adversarial, stochastic environments. Settlers of Catan prov…
Self-Resource Allocation in Multi-Agent LLM Systems
Alfonso Amayuelas, Jingbo Yang, Saaket Agashe +4
With the development of LLMs as agents, there is a growing interest in connecting multiple agents into multi-agent systems to solve tasks concurrently, focusing on their role in ta…