#benchmarking
85 papers · 1 filter
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
Wanyu Zhao, Wanbing Zhao
The paper studies whether large language model agents can discover statistical‑mechanical mappings for physics problems, introducing a benchmark of Ising‑type tasks and evaluating…
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Zihan Chen, Di Zhu, Lei Nico Zheng
The paper evaluates when large language models can reliably simulate human survey responses, finding systematic failures in individual-level predictions and demographic bias across…
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
Quim Motger, Marc Oriol, Jordi Marco +1
The paper surveys research on multi-agent debate for large language model systems, introduces a three‑dimensional taxonomy of participants, interaction mechanisms, and agreement pr…
Progress in Benchmarking Generics for Mathematical Computation
Daniel Pang, Stephen M. Watt
The paper presents SciGMark 1.5, a benchmark comparing specialized and generic implementations of numerical and symbolic kernels across modern languages, and analyzes how different…
ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science
Yuan Zhu, Ethan B. Liu, Frank Nie +1
ClinLens introduces a benchmark of 200 executable tasks that link multiple MIMIC data modalities—structured records, notes, ECGs, chest X‑rays, and echocardiograms—to evaluate long…
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models
Malena Loza, David Chushig-Muzo, Eva Milara +3
The paper empirically evaluates how nine tabular foundation models perform under various out-of-distribution shifts using real-world datasets, finding systematic performance degrad…