Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
RSNA Large Language Model Benchmark Dataset for Chest Radiographs of Cardiothoracic Disease: Radiologist Evaluation and Validation Enhanced by AI Labels (REVEAL-CXR)
Yishu Wei, Adam E. Flanders, Errol Colak +35
Multimodal large language models have demonstrated comparable performance to that of radiology trainees on multiple-choice board-style exams. However, to develop clinically useful…
cs.CL2025
A Multi-agent Large Language Model Framework to Automatically Assess Performance of a Clinical AI Triage Tool
Adam E. Flanders, Yifan Peng, Luciano Prevedello +6
Purpose: The purpose of this study was to determine if an ensemble of multiple LLM agents could be used collectively to provide a more reliable assessment of a pixel-based AI triag…
cs.CL2025
Generative Large Language Models Trained for Detecting Errors in Radiology Reports
Cong Sun, Kurt Teichman, Yiliang Zhou +8
In this retrospective study, a dataset was constructed with two parts. The first part included 1,656 synthetic chest radiology reports generated by GPT-4 using specified prompts, w…