#evaluation
19 resultsIFHierBench: Hierarchical Instruction Following for Large Language Models
Yuetian Mao, Chunyang Chen
The paper introduces IFHierBench, a benchmark for evaluating how well large language models follow hierarchical, nested constraints in instruction prompts, and shows current models…
Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
Fouad Bousetouane
The paper proposes the ProofAgent Index (PAI), a governance readiness framework for AI agents that assesses evaluation, context, compliance, and governance to determine production…
MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek
Ioakeim Perros, Cleopatra Papadopoulou, Ayoub Kirouane +1
The paper presents MORFES, a benchmark of 500 expert‑verified items for testing Greek language models' ability to recognize and generate inflected word forms, especially for low‑fr…
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
Jiajia Lin, Mingxuan Du, Tuowen Zhou +2
The paper presents MPIE-Bench, a benchmark of 2,500 multi-person interaction editing examples, and MPIE-Eval, an evaluation method that uses mesh reconstruction to assess anatomica…
Repair as Representational Work: Integration Bottlenecks in AI-Assisted Development
Daisaku Sato
The paper identifies an "integration bottleneck" in AI‑assisted programming where helpful suggestions reach the developer but cannot be turned into actionable instructions, and it…
ExplainBench: Evaluating Code Explanations from Agents
Zhiyuan Pan, Sungmin Kang, Imam Nur Bani Yusuf +1
The paper introduces ExplainBench, a benchmark that automatically evaluates how trustworthy the explanations generated by code‑writing LLM agents are, by checking if the explanatio…
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…
VAmoS Bench: Voice Agent Simulation Bench
Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5
The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…
Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
Zheng Tong, Yang Liu, Wanshu Fan +6
The paper reviews how large language and multimodal models are being used as autonomous agents in medical tasks, covering their architectures, applications, evaluation methods, and…
Idea2Plan: Exploring AI-Powered Research Planning
Jin Huang, Silviu Cucerzan, Sujay Kumar Jauhar +1
The paper studies how large language models can turn research ideas into detailed research plans, introducing the Idea2Plan benchmark and evaluating models like GPT‑5 on this task.
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Qiucheng Yu, Ruijie Xu, Mingang Chen +2
The paper introduces TSHA, a large benchmark of real-world indoor safety hazard assessment questions for evaluating vision‑language models, and shows that training on this data imp…
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
Seungone Kim, Ian Wu, Jinu Lee +8
The paper studies how using larger, chain‑of‑thought reasoning language models as evaluators—by allocating more test‑time compute—can improve the accuracy of evaluating and reranki…
How Inference Compute Shapes Frontier LLM Evaluation
Jessica McFadyen, Ole Jorgensen, Harry Coppock +2
The paper studies how the amount of compute allocated during inference (e.g., token budget, repeated attempts) affects the performance of frontier large language models on challeng…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
BuÄra Alperen Uluırmak, Rifat Kurban
The paper surveys recent work on evaluating large language models (LLMs) for safety and introduces the EvalSafetyGap framework to compare evaluation and alignment failures, illustr…
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
Faizan Faisal
The paper evaluates whether reasoning abilities of large language models improve the generation of structured clinical SOAP notes, finding that reasoning can actually degrade perfo…
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
Huatao Li, Xinwei Geng, Yuheng Wang +9
The paper presents DevicesWorld, a large executable benchmark of 6,140 tasks that require LLM‑based agents to operate across mobile, desktop, and IoT devices, and shows that curren…
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
Xi Fang, Weijie Xu, Yingqiang Ge +3
The paper introduces DRIFTLENS, a framework for measuring how injecting user-specific memory into personalized language models changes the models' reasoning steps, and evaluates me…
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen +4
The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…
Good Benchmarks
Ivan Bercovich
The paper outlines what makes a good benchmark task for AI, emphasizing that tasks should be correct, solvable, verifiable, well-specified, and challenging for meaningful reasons,…