#benchmarking

topicbenchmarking

85 papers · 1 filter

cs.AI2026

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Wanyu Zhao, Wanbing Zhao

The paper studies whether large language model agents can discover statistical‑mechanical mappings for physics problems, introducing a benchmark of Ising‑type tasks and evaluating…

cs.CL2026

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

Zihan Chen, Di Zhu, Lei Nico Zheng

The paper evaluates when large language models can reliably simulate human survey responses, finding systematic failures in individual-level predictions and demographic bias across…

cs.SE2026

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

Quim Motger, Marc Oriol, Jordi Marco +1

The paper surveys research on multi-agent debate for large language model systems, introduces a three‑dimensional taxonomy of participants, interaction mechanisms, and agreement pr…

cs.PL2026

Progress in Benchmarking Generics for Mathematical Computation

Daniel Pang, Stephen M. Watt

The paper presents SciGMark 1.5, a benchmark comparing specialized and generic implementations of numerical and symbolic kernels across modern languages, and analyzes how different…

cs.AI2026

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

Yuan Zhu, Ethan B. Liu, Frank Nie +1

ClinLens introduces a benchmark of 200 executable tasks that link multiple MIMIC data modalities—structured records, notes, ECGs, chest X‑rays, and echocardiograms—to evaluate long…

cs.LG2026

Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

Malena Loza, David Chushig-Muzo, Eva Milara +3

The paper empirically evaluates how nine tabular foundation models perform under various out-of-distribution shifts using real-world datasets, finding systematic performance degrad…