3 papers
cs.AI2026
Roundtable Policy: Confidence-Weighted-Consensus Aggregation Improves Multi-Agent-System Reasoning
Yu Yao, Jiayi Dong, Yang Yang +2
Multi-agent systems have demonstrated exceptional performance in downstream tasks beyond diverse single agent baselines. A growing body of work has explored ways to improve their r…
cs.AI2026
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
Jia-Kai Dong, I-Wei Huang, Chun-Tin Wu +1
We introduce ETOM, a five-level benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents within a hierarchical Model-Context Protocol (MCP) ecosystem. Existi…
cs.AI2026
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
Xinyi Wu, Jiagui Chen, Geng Hong +4
Web Agents are increasingly deployed to perform complex tasks in real web environments, yet their security evaluation remains fragmented and difficult to standardize. We present We…