From the 1 of 4 linked papers with an AI index.
4 papers
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
Shufan Jiang, Chios Chen, Zhiyang Chen
The autonomous discovery of bugs remains a significant challenge in modern software development. Compared to code generation, the complexity of dynamic runtime environments makes b…
NERdME: a Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories
Genet Asefa Gesese, Zongxiong Chen, Shufan Jiang +4
Existing scholarly information extraction (SIE) datasets focus on scientific papers and overlook implementation-level details in code repositories. README files describe datasets,…
NFDI4DS Shared Tasks for Scholarly Document Processing
Raia Abu Ahmad, Rana Abdulla, Tilahun Abedissa Taffa +18
Shared tasks are powerful tools for advancing research through community-based standardised evaluation. As such, they play a key role in promoting findable, accessible, interoperab…