3 papers
cs.CL2026
MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
Iker De la Iglesia, Johanna Ramirez-Romero, Jose Maria Villa-Gonzalez +3
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct d…
cs.CL2024
Ranking Over Scoring: Towards Reliable and Robust Automated Evaluation of LLM-Generated Medical Explanatory Arguments
Iker De la Iglesia, Iakes Goenaga, Johanna Ramirez-Romero +3
Evaluating LLM-generated text has become a key challenge, especially in domain-specific contexts like the medical field. This work introduces a novel evaluation methodology for LLM…
cs.CL2024
Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
Iker GarcÃa-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar +10
Research on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large langu…