NewEvery arXiv paper, its researchers & institutions — mapped.
the archive

#benchmark

86 results
cs.CV2026

EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits

Kaifan Zhang, Lihuo He, Yuqi Ji +4

The paper presents EEG-EditBench, a benchmark that uses controlled image edits to evaluate how EEG-to-image retrieval models capture fine-grained visual information.

#eeg decoding#image retrieval#visual perception#benchmark
cs.AI2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng, Yian Wang +20

The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…

#computer-use agents#vision-language models#reward modeling#benchmark
cs.CL2026

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5

The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…

#visual illusions#vision-language models#reasoning evaluation#benchmark
cs.AI2026

BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints

Ruslan Khrulev

The paper introduces BlueprintRepair, an interface that lets large language models make typed, local edits to Lean proof blueprints, and evaluates it on a benchmark of controlled p…

#theorem proving#lean#large language models#proof repair
cs.CL2026

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Tao Wen, Shuai Shao, Pei Ke +7

The paper introduces MiGUE-Bench, a benchmark that evaluates large language models on multi‑granularity event analysis tasks ranging from single‑document event detection to cross‑d…

#event detection#cross-document analysis#large language models#benchmark
cs.AI2026

IFHierBench: Hierarchical Instruction Following for Large Language Models

Yuetian Mao, Chunyang Chen

The paper introduces IFHierBench, a benchmark for evaluating how well large language models follow hierarchical, nested constraints in instruction prompts, and shows current models…

#instruction following#hierarchical constraints#large language models#benchmark
cs.CL2026

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

Ioakeim Perros, Cleopatra Papadopoulou, Ayoub Kirouane +1

The paper presents MORFES, a benchmark of 500 expert‑verified items for testing Greek language models' ability to recognize and generate inflected word forms, especially for low‑fr…

#morphology#inflection#benchmark#greek language
cs.AI2026

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

Shiwei Gan, Lichen Wang, Xiao Liu +4

The paper introduces Sign Language Question Answering (SLQA), a task where models answer natural language questions about sign language videos, and provides two benchmark datasets…

#sign language understanding#question answering#video-language#multimodal reasoning
cs.CL2026

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

Zheng Wu, Chenhao Xue, Shijie Zheng +3

The paper identifies a "salience bias" in large language models where explicit but irrelevant details cause the models to overlook implicit commonsense knowledge, and shows that th…

#commonsense reasoning#large language models#salience bias#prompt engineering
cs.CV2026

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Jiajia Lin, Mingxuan Du, Tuowen Zhou +2

The paper presents MPIE-Bench, a benchmark of 2,500 multi-person interaction editing examples, and MPIE-Eval, an evaluation method that uses mesh reconstruction to assess anatomica…

#multi-person interaction#image editing#benchmark#mesh reconstruction
cs.CV2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong +11

The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…

#visual reasoning#geometric reasoning#jigsaw puzzles#vision-language models
cs.CL2026

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Albert Gong, Kyuseong Choi, Abhineet Agarwal +5

The paper presents ORCA-bench, a benchmark that evaluates large language model agents on on-call root cause analysis tasks using real telemetry data from a live microservice system…

#root cause analysis#llm agents#benchmark#observability
cs.AI2026

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Weihang Wang, Kainan Tu, Jielei Zhang +9

The paper presents MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes that evaluates how large vision‑language models handle cultural and background knowledge, an…

#vision-language models#memes#cultural knowledge#benchmark
cs.CV2026

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

Yuyun Chen, Tianao Li, TianQuan Feng +4

The paper introduces EgoSafe-Bench, a first‑person video dataset and evaluation protocol designed to test causal and forensic reasoning for visual safety understanding, highlightin…

#egocentric video#visual safety#causal reasoning#benchmark
cs.CR2026

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Lehan Wang, Boli Chen, Ruixue Ding +7

The paper presents SecRespond, a benchmark that evaluates large language model agents on post-compromise incident‑response tasks using forensic disk snapshots, alerts, and vulnerab…

#incident response#large language models#benchmark#post-compromise
cs.CL2026

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

Arnav Hiray, Agam Shah, Caleb Lu +3

The paper presents CreditCardQA, a benchmark of 1,800 real‑world credit‑card agreement questions for testing numerical reasoning in language models, and shows that Program‑of‑Thoug…

#financial literacy#numerical reasoning#benchmark#chain-of-thought prompting
cs.CL2026

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

Jinhu Qi, Wentao Zhang, Siu Man Ng +4

The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…

#travel planning#llm agents#benchmark#constraint reasoning
cs.CL2026

APEX-Accounting

Julien Benchek, Austin Bennett, Jasmin Kern +8

The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…

#accounting automation#benchmark#large language models#evaluation metrics
cs.SD2026

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Weijie Wu, Junbo Li, Lin Li +2

The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…

#audio captioning#benchmark#evaluation metrics#large language models
cs.CL2026

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Ruxi Gu, Zhenliang Zhang, Wei Wang

The paper introduces ForgetBench, a benchmark for measuring how large language models retain or forget factual and relational knowledge when they are continuously edited over time.

#language models#continual learning#knowledge editing#benchmark
cs.AI2026

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Kawai Chung, Chunkit Chan, Yauwai Yim +12

The paper introduces MultivationBench, a benchmark that tests multimodal large language models on their ability to reason about evolving human motivations across sequential visual…

#multimodal reasoning#motivation inference#sequential reasoning#visual narratives
cs.CL2026

Benchmarking LLM Competence on Logical Inference over Probability Operators

Nayera Hasan, Jack Greff, Alvin Grissom

The paper presents a benchmark for testing large language models' ability to reason logically about probability expressions in English, and evaluates 29 models, revealing widesprea…

#logical reasoning#probability language#benchmark#bias analysis
cs.AI2026

VAmoS Bench: Voice Agent Simulation Bench

Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5

The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…

#voice agents#benchmark#simulation#customer support
cs.LG2026

PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective

Shengtian Yang, Yewen Li, Peng Jiang +4

The paper introduces PlatformBid, a benchmark for evaluating auto-bidding algorithms from the perspective of a unified advertising platform that combines SSP, DSP, and ad exchange…

#real-time bidding#auto-bidding#advertising platforms#benchmark
← prev1 / 4next →