1 paper
Iker De la Iglesia, Johanna Ramirez-Romero, Jose Maria Villa-Gonzalez +3
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct d…