From the 2 of 11 linked papers with an AI index.
11 papers
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
Hao Liang, Meiyi Qiang, Sizhe Qiu +2
The paper introduces WorkSurface-Bench, a benchmark that tests enterprise agents' ability to select the correct knowledge source (documents, tables, or graphs) before answering que…
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Runming He, Zhen Hao Wong, Hao Liang +4
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as per…
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Chengyu Shen, Yujie Fu, Gangtao Xin +13
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs
Zihan Nan, Yang Gu, Wei Liu +4
Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primaril…
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
Hao Liang, Qifeng Cai, Yibo Lin +11
The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-c…
VABench: A Comprehensive Benchmark for Audio-Video Generation
Daili Hua, Xizhi Wang, Bohan Zeng +6
Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks…