3 papers
cs.AI2026
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
Guo Chen, Ziwen Li, Reed Li +4
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across me…
cs.CL2026
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
Jiani Luo, Xiaoyan Zhao, Yang Zhang +4
Personalized dialogue requires more than recalling explicit user histories: systems also need to infer hidden user states that evolve through interaction and shape appropriate resp…
cs.AI2026
Towards Knowledgeable Deep Research: Framework and Benchmark
Wenxuan Liu, Zixuan Li, Long Bai +13
Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing…