4 papers
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following
Jaeyun Lee, Junyoung Koh, Zeynel Tok +2
Multi-constraint instruction following requires verifying whether a response satisfies multiple individual requirements, yet LLM judges are often assessed only through overall-resp…
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer +32
Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant…
Towards Understanding Multimodal Fine-Tuning: Spatial Features
Lachin Naghashyar, Hunar Batra, Ashkan Khakzar +4
Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-…
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…