Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
Wenrui Liu, Zixiang Liu, Elsie Dai +5
Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future tren…
cs.AI2025
SEDM: Scalable Self-Evolving Distributed Memory for Agents
Haoran Xu, Jiacong Hu, Ke Zhang +6
Long-term multi-agent systems inevitably generate vast amounts of trajectories and historical interactions, which makes efficient memory management essential for both performance a…