#evaluation

19 results
cs.AI2026

IFHierBench: Hierarchical Instruction Following for Large Language Models

Yuetian Mao, Chunyang Chen

The paper introduces IFHierBench, a benchmark for evaluating how well large language models follow hierarchical, nested constraints in instruction prompts, and shows current models…

#instruction following#hierarchical constraints#large language models#benchmark
cs.MA2026

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness

Fouad Bousetouane

The paper proposes the ProofAgent Index (PAI), a governance readiness framework for AI agents that assesses evaluation, context, compliance, and governance to determine production…

#ai agents#production readiness#governance#evaluation
cs.CL2026

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

Ioakeim Perros, Cleopatra Papadopoulou, Ayoub Kirouane +1

The paper presents MORFES, a benchmark of 500 expert‑verified items for testing Greek language models' ability to recognize and generate inflected word forms, especially for low‑fr…

#morphology#inflection#benchmark#greek language
cs.CV2026

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Jiajia Lin, Mingxuan Du, Tuowen Zhou +2

The paper presents MPIE-Bench, a benchmark of 2,500 multi-person interaction editing examples, and MPIE-Eval, an evaluation method that uses mesh reconstruction to assess anatomica…

#multi-person interaction#image editing#benchmark#mesh reconstruction
cs.HC2026

Repair as Representational Work: Integration Bottlenecks in AI-Assisted Development

Daisaku Sato

The paper identifies an "integration bottleneck" in AI‑assisted programming where helpful suggestions reach the developer but cannot be turned into actionable instructions, and it…

#ai-assisted programming#integration bottleneck#repair#evaluation
cs.SE2026

ExplainBench: Evaluating Code Explanations from Agents

Zhiyuan Pan, Sungmin Kang, Imam Nur Bani Yusuf +1

The paper introduces ExplainBench, a benchmark that automatically evaluates how trustworthy the explanations generated by code‑writing LLM agents are, by checking if the explanatio…

#code explanations#large language model agents#benchmarking#evaluation
cs.CL2026

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

Jinhu Qi, Wentao Zhang, Siu Man Ng +4

The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…

#travel planning#llm agents#benchmark#constraint reasoning
cs.AI2026

VAmoS Bench: Voice Agent Simulation Bench

Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5

The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…

#voice agents#benchmark#simulation#customer support
cs.CV2026

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

Zheng Tong, Yang Liu, Wanshu Fan +6

The paper reviews how large language and multimodal models are being used as autonomous agents in medical tasks, covering their architectures, applications, evaluation methods, and…

#agentic ai#clinical decision support#multimodal models#evaluation
cs.CL2026

Idea2Plan: Exploring AI-Powered Research Planning

Jin Huang, Silviu Cucerzan, Sujay Kumar Jauhar +1

The paper studies how large language models can turn research ideas into detailed research plans, introducing the Idea2Plan benchmark and evaluating models like GPT‑5 on this task.

#research planning#large language models#benchmarking#autonomous research agents
cs.CV2026

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

Qiucheng Yu, Ruijie Xu, Mingang Chen +2

The paper introduces TSHA, a large benchmark of real-world indoor safety hazard assessment questions for evaluating vision‑language models, and shows that training on this data imp…

#vision-language models#safety assessment#benchmark#indoor hazards
cs.CL2026

Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Seungone Kim, Ian Wu, Jinu Lee +8

The paper studies how using larger, chain‑of‑thought reasoning language models as evaluators—by allocating more test‑time compute—can improve the accuracy of evaluating and reranki…

#evaluation#reasoning models#test-time compute#chain-of-thought
cs.AI2026

How Inference Compute Shapes Frontier LLM Evaluation

Jessica McFadyen, Ole Jorgensen, Harry Coppock +2

The paper studies how the amount of compute allocated during inference (e.g., token budget, repeated attempts) affects the performance of frontier large language models on challeng…

#large language models#evaluation#inference compute#benchmark scaling
cs.AI2026

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

Buğra Alperen Uluırmak, Rifat Kurban

The paper surveys recent work on evaluating large language models (LLMs) for safety and introduces the EvalSafetyGap framework to compare evaluation and alignment failures, illustr…

#large language models#evaluation#ai safety#benchmarking
cs.CL2026

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

Faizan Faisal

The paper evaluates whether reasoning abilities of large language models improve the generation of structured clinical SOAP notes, finding that reasoning can actually degrade perfo…

#clinical note generation#large language models#reasoning#retrieval-augmented generation
cs.CL2026

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

Huatao Li, Xinwei Geng, Yuheng Wang +9

The paper presents DevicesWorld, a large executable benchmark of 6,140 tasks that require LLM‑based agents to operate across mobile, desktop, and IoT devices, and shows that curren…

#cross-device interaction#llm agents#benchmark#multimodal environments
cs.AI2026

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

Xi Fang, Weijie Xu, Yingqiang Ge +3

The paper introduces DRIFTLENS, a framework for measuring how injecting user-specific memory into personalized language models changes the models' reasoning steps, and evaluates me…

#personalization#language models#reasoning drift#memory injection
cs.CL2026

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Peiyu Li, Xiuxiu Tang, Si Chen +4

The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…

#large language models#evaluation#adaptive testing#item response theory
cs.AI2026

Good Benchmarks

Ivan Bercovich

The paper outlines what makes a good benchmark task for AI, emphasizing that tasks should be correct, solvable, verifiable, well-specified, and challenging for meaningful reasons,…

#benchmarks#evaluation#task design#dataset quality