3 papers
cs.CL2026
From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations
Ping Wang, Xiangguo Sun, Bingbing Xu +2
Large language models (LLMs) frequently produce confident yet factually incorrect responses when user inputs contain misleading premises, a phenomenon we attribute to fact perturba…
cs.AI2026
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
Guo Chen, Ziwen Li, Reed Li +4
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across me…
cs.AI2026
Towards Knowledgeable Deep Research: Framework and Benchmark
Wenxuan Liu, Zixuan Li, Long Bai +13
Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing…