2 citations · 2 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont +10
Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical comput…
cs.AI2026
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Boyan Li, Zhuowen Liang, Yupeng Xie +11
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multi…