works on

From the 2 of 19 linked papers with an AI index.

collaborators

19 papers

cs.CL2026

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

Hao Liang, Meiyi Qiang, Sizhe Qiu +2

The paper introduces WorkSurface-Bench, a benchmark that tests enterprise agents' ability to select the correct knowledge source (documents, tables, or graphs) before answering que…

cs.CL2026

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Chengyu Shen, Yujie Fu, Gangtao Xin +13

Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…

cs.CL2026

GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs

Hengyi Feng, Zeang Sheng, Meiyi Qiang +2

Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenarios, such as citation network…

cs.CV2026

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

DataFlow Team, Bohan Zeng, Daili Hua +39

World models have garnered significant attention as a promising research direction in artificial intelligence, yet a clear and unified definition remains lacking. In this paper, we…

cs.LG2026

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Hao Liang, Qifeng Cai, Yibo Lin +11

The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-c…

cs.MM2026

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

Qijie You, Hao Liang, Mingrui Chen +4

As video becomes increasingly central to information dissemination and multimodal large language models (MLLMs) continue to advance, evaluating video retrieval has become increasin…