#benchmarking

85 results
cs.AI2026

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness

Debin Meng, Jiaming Yang, Zefang Zong +4

The paper introduces DataClawEval, a benchmark that tests autonomous LLM agents on end-to-end data engineering tasks across multiple production-grade SQL and Spark engines using de…

#data engineering#large language models#autonomous agents#benchmarking
astro-ph.IM2026

DB-Bench: Benchmarking Deblenders for LSST DESC Using the Blending ToolKit

Aidan Berres, Grant Merz, Xin Liu +3

The paper evaluates several image deblending algorithms for LSST galaxy surveys using the Blending ToolKit, comparing their detection, segmentation, and reconstruction performance…

#deblending#galaxy surveys#lsst#image processing
cs.AI2026

InfoOps Bench: A live information operations safety benchmark

Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright +1

The paper introduces InfoOps Bench, a continuously updated benchmark that measures how easily frontier language models can be co-opted for state-backed information operations, usin…

#information operations#model safety#benchmarking#language models
cs.CL2026

FinanceHarness: Autonomous Financial Deep Research Framework

Yijia Xiao, Rujun Han, Yanfei Chen +8

The paper introduces FinanceHarness, a framework that uses large language models and autonomous agents to automate end‑to‑end financial deep research, and presents FinanceGym, a be…

#financial research automation#large language models#autonomous agents#benchmarking
cs.CV2026

Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation

Jinghong Liu, Yuchuan Deng, Fanping Liu +2

The paper presents FAME, a unified benchmark for evaluating few-shot medical image segmentation methods across multiple anatomical sites, imaging modalities, and settings, and anal…

#few-shot segmentation#medical imaging#benchmarking#zero-shot learning
cs.AI2026

An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding

Yanyu Ren, Yunfeng Bai, Xizheng Wang +2

The paper presents MSEval, a benchmark that evaluates how multi‑agent coding systems build real‑world software, measuring functional success, latency, and token cost while varying…

#multi-agent coding#software development#benchmarking#coordination topology
quant-ph2026

Benchmarking Quantum Simulations of the Lipkin-Meshkov-Glick Model Using Large Tensor Networks

Maggie Bao, Rushil Dandamudi, Jerimiah Wright +4

The paper benchmarks classical tensor‑network methods (DMRG) against NISQ quantum algorithms (VQE and SQD) for computing ground‑state energies of the Lipkin‑Meshkov‑Glick model, pr…

#quantum simulation#tensor networks#dmrg#vqe
cs.AI2026

Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

Phuc Pham, Truong-Son Hy

The paper proposes a Function–Evidence–Validation (FEV) framework to evaluate bioinformatics workflows generated by large language model agents, emphasizing workflow correctness an…

#bioinformatics#large language models#agentic systems#workflow evaluation
cs.LG2026

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Xingjian Wu, Xuhang Zhu, Xingchen Liu +6

The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…

#large language models#autonomous agents#benchmarking#process evaluation
cs.CR2026

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense

Gal Engelberg, Michael Arenzon, Leon Goldberg

The paper introduces the Open Security Benchmark (OSB), a framework that provides a frozen, holistic enterprise security dataset and evaluation tools for testing autonomous AI agen…

#autonomous cyber defense#security posture management#benchmarking#enterprise security
cs.CV2026

VETO: Towards Protecting Images From Frontier AI Editing

Jonas Grebe, Hossein Shakibania, Tobias Braun +2

The paper presents VETO, a subtle anti-edit cloak that disrupts how modern diffusion-based image editors read source images, and introduces VetoBench, a benchmark for evaluating pr…

#image editing protection#diffusion models#joint attention#anti-edit defenses
cs.CV2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

Siyao Li, Li Siyao, Jiawei Gu +16

The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…

#vision-language models#embodied AI#benchmarking#physical simulation
cs.CL2026

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

Linyu Li, Zhi Jin, Yichi Zhang +6

The paper introduces EC-Reason-Bench, a training-free diagnostic benchmark that evaluates why general large language models struggle with detailed enzyme classification and how per…

#enzyme classification#large language models#zero-shot evaluation#open-book reasoning
cs.SE2026

ExplainBench: Evaluating Code Explanations from Agents

Zhiyuan Pan, Sungmin Kang, Imam Nur Bani Yusuf +1

The paper introduces ExplainBench, a benchmark that automatically evaluates how trustworthy the explanations generated by code‑writing LLM agents are, by checking if the explanatio…

#code explanations#large language model agents#benchmarking#evaluation
cs.AI2026

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Wanyu Zhao, Wanbing Zhao

The paper studies whether large language model agents can discover statistical‑mechanical mappings for physics problems, introducing a benchmark of Ising‑type tasks and evaluating…

#large language models#statistical mechanics#physics problem solving#benchmarking
cs.LG2026

Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts

Keith G. Mills, Aedan J. DeFrates, Joong Ho Kim

The paper evaluates how different graph neural network message‑passing layers perform on scalar regression tasks, finding that deep convolutional GNNs like GEN generally outperform…

#graph neural networks#message passing#regression#benchmarking
cs.SD2026

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Jiachen Qian, Junyu Li

The paper investigates how variations in speech prosody, while keeping the transcript unchanged, can cause jailbreaks in audio-capable language models, introducing a new evaluation…

#audio language models#jailbreak attacks#prosody manipulation#safety evaluation
cs.DC2026

Mind the Gap: The Disconnect Between Synthetic and Natural Edge Weights in Parallel Single-Source Shortest Path

Marco D'Antonio, Thai Son Mai, Hans Vandierendonck

The paper investigates how synthetic uniform edge weight distributions used in benchmarks differ from the heavy‑tailed distributions found in real graphs, showing that parallel sin…

#single-source shortest path#edge weight distribution#benchmarking#parallel algorithms
cs.RO2026

NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and Education

Koneshka Bandyopadhyay, Ansh Mehta, Bassel El Mabsout +1

The paper introduces NeoRacer, an open-source 1:12 scale autonomous race car equipped with high-performance compute and sensors, designed as an affordable, standardized platform fo…

#autonomous racing#scale models#open-source hardware#benchmarking
cs.AI2026

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Jingbo Zhou, Yusai Zhao, Qi Bao +12

The paper presents OmegaUse-OfficeVal, a benchmark that evaluates large language model agents on long‑horizon office‑suite tasks while providing economic signals (human labor time…

#llm agents#office suite tasks#benchmarking#economic evaluation
cs.LG2026

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance

Chao Peng, Zhiheng Lyu, Peijie Dong +2

The paper proposes a benchmark metric called the horizon residual to compare long-horizon task success against predictions from short-stage baselines, highlighting how performance…

#long-horizon evaluation#benchmarking#agent performance#context rot
cs.SE2026

AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents

Jiajun Zhou, Zhaoxuan Ke, Jihang Ye +3

The paper presents AgentS4D, a sandboxed benchmark that evaluates runtime safety risks of large language model‑based workspace agents throughout their execution lifecycle, using a…

#llm agents#runtime safety#benchmarking#risk assessment
cs.SE2026

SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements

Pengyu Xue, He Yang Yuan, Xin Wang +6

The paper introduces SWE-NFI, a benchmark that assesses how coding agents can make non-functional, behavior-preserving improvements to Python code, using real pull‑request tasks an…

#coding agents#non-functional improvements#benchmarking#software quality
cs.AI2026

PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments

Haoyu Chen, Xirui Shi, Yuyao Wang +2

PAUSE is a benchmark that evaluates personal AI assistants on their ability to manage persistent user state, respect configurations and permissions, and coordinate actions across m…

#personal assistants#benchmarking#stateful reasoning#service integration
← Prev1 / 4Next →