From the 2 of 25 linked papers with an AI index.
23 citations · 23 across the 10 of their papers we have counts for
11 papers · 1 filter
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer +32
Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant…
Improving the Performance of Radiology Report De-identification with Large-Scale Training and Benchmarking Against Cloud Vendor Methods
Eva Prakash, Maayane Attias, Pierre Chambon +5
Objective: To enhance automated de-identification of radiology reports by scaling transformer-based models through extensive training datasets and benchmarking performance against…
Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
Johannes Moll, Markus Graf, Tristan Lemke +7
Vision-language models (VLMs) often produce chain-of-thought (CoT) explanations that sound plausible yet fail to reflect the underlying decision process, undermining trust in high-…
Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports
Alois Thomas, Maya Varma, Jean-Benoit Delbrouck +1
Automating radiology report generation with Large Vision-Language Models (LVLMs) holds great potential, yet these models often produce clinically critical hallucinations, posing se…
RadEval: A framework for radiology text evaluation
Justin Xu, Xi Zhang, Javid Abderezaei +9
We introduce RadEval, a unified, open-source framework for evaluating radiology texts. RadEval consolidates a diverse range of metrics, from classic n-gram overlap (BLEU, ROUGE) an…
Structuring Radiology Reports: Challenging LLMs with Lightweight Models
Johannes Moll, Louisa Fay, Asfandyar Azhar +5
Radiology reports are critical for clinical decision-making but often lack a standardized format, limiting both human interpretability and machine learning (ML) applications. While…