2 papers
cs.CL2026
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang +57
Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation. As the field enters the "…
cs.CL2024
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
Wei Wang, Dan Zhang, Tao Feng +2
Large Language Models (LLMs) are becoming increasingly powerful and capable of handling complex tasks, e.g., building single agents and multi-agent systems. Compared to single agen…