4 papers · 1 filter
From Attribution to Abstention: Training-Free Attention-Based Auditing for Clinical Summarization
Qianqi Yan, Huy Nguyen, Sumana Srivatsa +3
Deploying multimodal large language models (MLLMs) for clinical summarization demands not only fluent generation but also transparency about where each statement originates-and a m…
GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
Xiuyuan Chen, Tao Sun, Dexin Su +37
Current benchmarks for AI clinician systems, often based on multiple-choice exams or manual rubrics, fail to capture the depth, robustness, and safety required for real-world clini…
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
Seokhee Hong, Sunkyoung Kim, Guijin Son +3
The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applica…
Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models
Zongru Shao, Xin Wang, Zhanyang Liu +2
Recent research leverages large language models (LLMs) for early mental health detection, such as depression, often optimized with machine-generated data. However, their detection…