Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
Guo Chen, Ziwen Li, Reed Li +4
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across me…
cs.AI2026
Towards Knowledgeable Deep Research: Framework and Benchmark
Wenxuan Liu, Zixuan Li, Long Bai +13
Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing…