#benchmark
86 resultsEEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits
Kaifan Zhang, Lihuo He, Yuqi Ji +4
The paper presents EEG-EditBench, a benchmark that uses controlled image edits to evaluate how EEG-to-image retrieval models capture fine-grained visual information.
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5
The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…
BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints
Ruslan Khrulev
The paper introduces BlueprintRepair, an interface that lets large language models make typed, local edits to Lean proof blueprints, and evaluates it on a benchmark of controlled p…
From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
Tao Wen, Shuai Shao, Pei Ke +7
The paper introduces MiGUE-Bench, a benchmark that evaluates large language models on multi‑granularity event analysis tasks ranging from single‑document event detection to cross‑d…
IFHierBench: Hierarchical Instruction Following for Large Language Models
Yuetian Mao, Chunyang Chen
The paper introduces IFHierBench, a benchmark for evaluating how well large language models follow hierarchical, nested constraints in instruction prompts, and shows current models…
MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek
Ioakeim Perros, Cleopatra Papadopoulou, Ayoub Kirouane +1
The paper presents MORFES, a benchmark of 500 expert‑verified items for testing Greek language models' ability to recognize and generate inflected word forms, especially for low‑fr…
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
Shiwei Gan, Lichen Wang, Xiao Liu +4
The paper introduces Sign Language Question Answering (SLQA), a task where models answer natural language questions about sign language videos, and provides two benchmark datasets…
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
Zheng Wu, Chenhao Xue, Shijie Zheng +3
The paper identifies a "salience bias" in large language models where explicit but irrelevant details cause the models to overlook implicit commonsense knowledge, and shows that th…
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
Jiajia Lin, Mingxuan Du, Tuowen Zhou +2
The paper presents MPIE-Bench, a benchmark of 2,500 multi-person interaction editing examples, and MPIE-Eval, an evaluation method that uses mesh reconstruction to assess anatomica…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Albert Gong, Kyuseong Choi, Abhineet Agarwal +5
The paper presents ORCA-bench, a benchmark that evaluates large language model agents on on-call root cause analysis tasks using real telemetry data from a live microservice system…
MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
Weihang Wang, Kainan Tu, Jielei Zhang +9
The paper presents MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes that evaluates how large vision‑language models handle cultural and background knowledge, an…
EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
Yuyun Chen, Tianao Li, TianQuan Feng +4
The paper introduces EgoSafe-Bench, a first‑person video dataset and evaluation protocol designed to test causal and forensic reasoning for visual safety understanding, highlightin…
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Lehan Wang, Boli Chen, Ruixue Ding +7
The paper presents SecRespond, a benchmark that evaluates large language model agents on post-compromise incident‑response tasks using forensic disk snapshots, alerts, and vulnerab…
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
Arnav Hiray, Agam Shah, Caleb Lu +3
The paper presents CreditCardQA, a benchmark of 1,800 real‑world credit‑card agreement questions for testing numerical reasoning in language models, and shows that Program‑of‑Thoug…
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern +8
The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Weijie Wu, Junbo Li, Lin Li +2
The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…
ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
Ruxi Gu, Zhenliang Zhang, Wei Wang
The paper introduces ForgetBench, a benchmark for measuring how large language models retain or forget factual and relational knowledge when they are continuously edited over time.
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Kawai Chung, Chunkit Chan, Yauwai Yim +12
The paper introduces MultivationBench, a benchmark that tests multimodal large language models on their ability to reason about evolving human motivations across sequential visual…
Benchmarking LLM Competence on Logical Inference over Probability Operators
Nayera Hasan, Jack Greff, Alvin Grissom
The paper presents a benchmark for testing large language models' ability to reason logically about probability expressions in English, and evaluates 29 models, revealing widesprea…
VAmoS Bench: Voice Agent Simulation Bench
Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5
The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…
PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective
Shengtian Yang, Yewen Li, Peng Jiang +4
The paper introduces PlatformBid, a benchmark for evaluating auto-bidding algorithms from the perspective of a unified advertising platform that combines SSP, DSP, and ad exchange…