works on

From the 2 of 14 linked papers with an AI index.

collaborators

14 papers

cs.CL2026

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

Hao Liang, Meiyi Qiang, Sizhe Qiu +2

The paper introduces WorkSurface-Bench, a benchmark that tests enterprise agents' ability to select the correct knowledge source (documents, tables, or graphs) before answering que…

cs.AI2026

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

Hao Jiang, Gangtao Xin, Yingdi Huang +35

Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…

cs.CL2026

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Chengyu Shen, Yujie Fu, Gangtao Xin +13

Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…

cs.DB2026

CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs

Zihan Nan, Yang Gu, Wei Liu +4

Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primaril…

cs.AI2026

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows

Wei Liu, Yang Gu, Xi Yan +5

Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines. While recent LLM-based approac…

cs.CV2026

VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction

Kaixin Zhu, Yiwen Tang, Yifan Yang +9

High-quality 3D scene reconstruction has recently advanced toward generalizable feed-forward architectures, enabling the generation of complex environments in a single forward pass…