benchmark dataset 1enterprise agents 1heterogeneous knowledge sources 1knowledge routing 1multi-surface retrieval 1tool use 1
From the 2 of 11 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
Hao Liang, Meiyi Qiang, Sizhe Qiu +2
The paper introduces WorkSurface-Bench, a benchmark that tests enterprise agents' ability to select the correct knowledge source (documents, tables, or graphs) before answering que…
cs.CL2026
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Chengyu Shen, Yujie Fu, Gangtao Xin +13
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…
cs.CL2026
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
Chengyu Shen, Yanheng Hou, Minghui Pan +8
Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify approp…