3 papers
cs.AI2026
An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding
Yanyu Ren, Yunfeng Bai, Xizheng Wang +2
The paper presents MSEval, a benchmark that evaluates how multi‑agent coding systems build real‑world software, measuring functional success, latency, and token cost while varying…
cs.AI2026
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
Ying Mo, Yu Bai, Dapeng Sun +4
Recent advances in Multimodal Large Language Models (MLLMs) have enabled agents to operate in open-ended web and operating system environments. However, existing benchmarks predomi…
cs.CR2025
The Digital Cybersecurity Expert: How Far Have We Come?
Dawei Wang, Geng Zhou, Xianglong Li +5
The increasing deployment of large language models (LLMs) in the cybersecurity domain underscores the need for effective model selection and evaluation. However, traditional evalua…