activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

Iker De la Iglesia, Johanna Ramirez-Romero, Jose Maria Villa-Gonzalez +3

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct d…

cs.CL2026

To Adapt or not to Adapt, Rethinking the Value of Medical Knowledge-Aware Large Language Models

Ane G. Domingo-Aldama, Iker De La Iglesia, Maitane Urruela +2

BACKGROUND: Recent studies have shown that domain-adapted large language models (LLMs) do not consistently outperform general-purpose counterparts on standard medical benchmarks, r…

cs.CL2025

ArgHiTZ at ArchEHR-QA 2025: A Two-Step Divide and Conquer Approach to Patient Question Answering for Top Factuality

Adrián Cuadrón, Aimar Sagasti, Maitane Urruela +5

This work presents three different approaches to address the ArchEHR-QA 2025 Shared Task on automated patient question answering. We introduce an end-to-end prompt-based baseline a…

cs.CL2025

Medical Argument Mining: Exploitation of Scarce Data Using NLI Systems

Maitane Urruela, Sergio Martín, Iker De la Iglesia +1

This work presents an Argument Mining process that extracts argumentative entities from clinical texts and identifies their relationships using token classification and Natural Lan…

cs.CL2024

Ranking Over Scoring: Towards Reliable and Robust Automated Evaluation of LLM-Generated Medical Explanatory Arguments

Iker De la Iglesia, Iakes Goenaga, Johanna Ramirez-Romero +3

Evaluating LLM-generated text has become a key challenge, especially in domain-specific contexts like the medical field. This work introduces a novel evaluation methodology for LLM…