10 papers
What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA
Yuelyu Ji, Min Gu Kwak, Hang Zhang +3
Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. \textbf{We find that the checker's \emph{output distribu…
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
Chenyu Li, Zohaib Akhtar, Mingu Kwak +13
As large language models (LLMs) increasingly generate and process clinical text, scalable evaluation has become critical. LLM-as-a-Judge (LaaJ), which uses LLMs to evaluate model o…
Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
Hang Zhang, Ruheng Wang, Yuelyu Ji +7
Large language models have achieved strong performance on medical reasoning benchmarks, yet their deployment in clinical settings demands rigorous verification to ensure factual ac…
MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
Yuelyu Ji, Min Gu Kwak, Hang Zhang +3
Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with…
Orchestrator Multi-Agent Clinical Decision Support System for Secondary Headache Diagnosis in Primary Care
Xizhi Wu, Nelly Estefanie Garduno-Rapp, Justin F Rousseau +6
Unlike most primary headaches, secondary headaches need specialized care and can have devastating consequences if not treated promptly. Clinical guidelines highlight several 'red f…
Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing
Xizhi Wu, Madeline S. Kreider, Philip E. Empey +2
Objective: Fluoropyrimidines are widely prescribed for colorectal and breast cancers, but are associated with toxicities such as hand-foot syndrome and cardiotoxicity. Since toxici…