3 papers
cs.CL2026
EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries
Jiyoun Kim, Muhan Yeo, Eunhye Jang +14
Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmissio…
cs.CV2026
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
Byungjin Choi, Seongsu Bae, Sunjun Kweon +1
We introduce KorMedMCQA-V, a Korean medical licensing-exam-style multimodal multiple-choice question answering benchmark for evaluating vision-language models (VLMs). The dataset c…
cs.CY2025
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
Sunjun Kweon, Sooyohn Nam, Hyunseung Lim +2
Virtual Teaching Assistants (VTAs) powered by Large Language Models (LLMs) have the potential to enhance student learning by providing instant feedback and facilitating multi-turn…