4 papers
ProtocolBench: Which LLM MultiAgent Protocol to Choose?
Hongyi Du, Jiaqi Su, Jisen Li +6
As large-scale multi-agent systems evolve, the communication protocol layer has become a critical yet under-evaluated factor shaping performance and reliability. Despite the existe…
SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
Huacan Chai, Yukai Wang, Yingxuan Yang +7
Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across in…
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
Chenyu Zhou, Huacan Chai, Wenteng Chen +18
Large language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the…
Where LLM Agents Fail and How They can Learn From Failures
Kunlun Zhu, Zijia Liu, Bingxuan Li +15
Large Language Model (LLM) agents, which integrate planning, memory, reflection, and tool-use modules, have shown promise in solving complex, multi-step tasks. Yet their sophistica…