Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
Iker De la Iglesia, Johanna Ramirez-Romero, Jose Maria Villa-Gonzalez +3
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct d…
cs.CL2026
To Adapt or not to Adapt, Rethinking the Value of Medical Knowledge-Aware Large Language Models
Ane G. Domingo-Aldama, Iker De La Iglesia, Maitane Urruela +2
BACKGROUND: Recent studies have shown that domain-adapted large language models (LLMs) do not consistently outperform general-purpose counterparts on standard medical benchmarks, r…
cs.CL2025
ArgHiTZ at ArchEHR-QA 2025: A Two-Step Divide and Conquer Approach to Patient Question Answering for Top Factuality
Adrián Cuadrón, Aimar Sagasti, Maitane Urruela +5
This work presents three different approaches to address the ArchEHR-QA 2025 Shared Task on automated patient question answering. We introduce an end-to-end prompt-based baseline a…