15 papers
Repo2Skill-Evo: Repository Skills Go Stale in Silence
Chenyuan Duan, Ge Shi, Zineng Mao +10
Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural knowledge: which APIs to call, w…
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs
Zihan Nan, Yang Gu, Wei Liu +4
Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primaril…
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
Hao Liang, Meiyi Qiang, Sizhe Qiu +2
Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships. Exis…
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
Hao Jiang, Gangtao Xin, Yingdi Huang +35
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Chengyu Shen, Yujie Fu, Gangtao Xin +13
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
Wei Liu, Yang Gu, Xi Yan +5
Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines. While recent LLM-based approac…