#benchmarking
85 resultsDataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness
Debin Meng, Jiaming Yang, Zefang Zong +4
The paper introduces DataClawEval, a benchmark that tests autonomous LLM agents on end-to-end data engineering tasks across multiple production-grade SQL and Spark engines using de…
DB-Bench: Benchmarking Deblenders for LSST DESC Using the Blending ToolKit
Aidan Berres, Grant Merz, Xin Liu +3
The paper evaluates several image deblending algorithms for LSST galaxy surveys using the Blending ToolKit, comparing their detection, segmentation, and reconstruction performance…
InfoOps Bench: A live information operations safety benchmark
Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright +1
The paper introduces InfoOps Bench, a continuously updated benchmark that measures how easily frontier language models can be co-opted for state-backed information operations, usin…
FinanceHarness: Autonomous Financial Deep Research Framework
Yijia Xiao, Rujun Han, Yanfei Chen +8
The paper introduces FinanceHarness, a framework that uses large language models and autonomous agents to automate end‑to‑end financial deep research, and presents FinanceGym, a be…
Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation
Jinghong Liu, Yuchuan Deng, Fanping Liu +2
The paper presents FAME, a unified benchmark for evaluating few-shot medical image segmentation methods across multiple anatomical sites, imaging modalities, and settings, and anal…
An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding
Yanyu Ren, Yunfeng Bai, Xizheng Wang +2
The paper presents MSEval, a benchmark that evaluates how multi‑agent coding systems build real‑world software, measuring functional success, latency, and token cost while varying…
Benchmarking Quantum Simulations of the Lipkin-Meshkov-Glick Model Using Large Tensor Networks
Maggie Bao, Rushil Dandamudi, Jerimiah Wright +4
The paper benchmarks classical tensor‑network methods (DMRG) against NISQ quantum algorithms (VQE and SQD) for computing ground‑state energies of the Lipkin‑Meshkov‑Glick model, pr…
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
Phuc Pham, Truong-Son Hy
The paper proposes a Function–Evidence–Validation (FEV) framework to evaluate bioinformatics workflows generated by large language model agents, emphasizing workflow correctness an…
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
Xingjian Wu, Xuhang Zhu, Xingchen Liu +6
The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…
Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
Gal Engelberg, Michael Arenzon, Leon Goldberg
The paper introduces the Open Security Benchmark (OSB), a framework that provides a frozen, holistic enterprise security dataset and evaluation tools for testing autonomous AI agen…
VETO: Towards Protecting Images From Frontier AI Editing
Jonas Grebe, Hossein Shakibania, Tobias Braun +2
The paper presents VETO, a subtle anti-edit cloak that disrupts how modern diffusion-based image editors read source images, and introduces VetoBench, a benchmark for evaluating pr…
HumanCLAW: Can Vision-Language Models Act Through a Body?
Siyao Li, Li Siyao, Jiawei Gu +16
The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…
Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
Linyu Li, Zhi Jin, Yichi Zhang +6
The paper introduces EC-Reason-Bench, a training-free diagnostic benchmark that evaluates why general large language models struggle with detailed enzyme classification and how per…
ExplainBench: Evaluating Code Explanations from Agents
Zhiyuan Pan, Sungmin Kang, Imam Nur Bani Yusuf +1
The paper introduces ExplainBench, a benchmark that automatically evaluates how trustworthy the explanations generated by code‑writing LLM agents are, by checking if the explanatio…
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
Wanyu Zhao, Wanbing Zhao
The paper studies whether large language model agents can discover statistical‑mechanical mappings for physics problems, introducing a benchmark of Ising‑type tasks and evaluating…
Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts
Keith G. Mills, Aedan J. DeFrates, Joong Ho Kim
The paper evaluates how different graph neural network message‑passing layers perform on scalar regression tasks, finding that deep convolutional GNNs like GEN generally outperform…
Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
Jiachen Qian, Junyu Li
The paper investigates how variations in speech prosody, while keeping the transcript unchanged, can cause jailbreaks in audio-capable language models, introducing a new evaluation…
Mind the Gap: The Disconnect Between Synthetic and Natural Edge Weights in Parallel Single-Source Shortest Path
Marco D'Antonio, Thai Son Mai, Hans Vandierendonck
The paper investigates how synthetic uniform edge weight distributions used in benchmarks differ from the heavy‑tailed distributions found in real graphs, showing that parallel sin…
NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and Education
Koneshka Bandyopadhyay, Ansh Mehta, Bassel El Mabsout +1
The paper introduces NeoRacer, an open-source 1:12 scale autonomous race car equipped with high-performance compute and sensors, designed as an affordable, standardized platform fo…
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Jingbo Zhou, Yusai Zhao, Qi Bao +12
The paper presents OmegaUse-OfficeVal, a benchmark that evaluates large language model agents on long‑horizon office‑suite tasks while providing economic signals (human labor time…
Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
Chao Peng, Zhiheng Lyu, Peijie Dong +2
The paper proposes a benchmark metric called the horizon residual to compare long-horizon task success against predictions from short-stage baselines, highlighting how performance…
AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents
Jiajun Zhou, Zhaoxuan Ke, Jihang Ye +3
The paper presents AgentS4D, a sandboxed benchmark that evaluates runtime safety risks of large language model‑based workspace agents throughout their execution lifecycle, using a…
SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements
Pengyu Xue, He Yang Yuan, Xin Wang +6
The paper introduces SWE-NFI, a benchmark that assesses how coding agents can make non-functional, behavior-preserving improvements to Python code, using real pull‑request tasks an…
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
Haoyu Chen, Xirui Shi, Yuyao Wang +2
PAUSE is a benchmark that evaluates personal AI assistants on their ability to manage persistent user state, respect configurations and permissions, and coordinate actions across m…