most citedDataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows

1 citations · 1 across the 16 of their papers we have counts for

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

ReproAgent: Contract-Guided Paper-to-Code Reproduction

Xue Hu, Zewei Pan, Zhongyuan Wang +3

Paper-to-code reproduction asks scientific AI agents to turn research papers into executable repositories that preserve the paper's method, protocol and artifacts. This is difficul…

cs.AI2026

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

Xue Hu, Zewei Pan, Zeli Su +2

LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations. We define this failure mode as semantic drift, where generated code si…

cs.AI2026

ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models

Zhou Liu, Chaoyang Han, Zewei Pan +2

Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prompt labels disconnected from learned behav…

cs.AI2026

SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation

Zhou Liu, Ligang Huang, Zeli Su +5

Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying whi…

cs.AI2026

DataCross: A Unified Benchmark and Agent Framework for Cross-Modal Heterogeneous Data Analysis

Ruyi Qi, Zhou Liu, Wentao Zhang

In real-world data science and enterprise decision-making, critical information is often fragmented across directly queryable structured sources (e.g., SQL, CSV) and "zombie data"…

cs.AI2025★ 1 cited

DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows

Zhou Liu, Zhaoyang Han, Guochen Yan +5

Data governance ensures data quality, security, and compliance through policies and standards, a critical foundation for scaling modern AI development. Recently, large language mod…