activity
20242026
collaborators

9 papers

cs.CY2026

A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework

Chenyu Li, Zohaib Akhtar, Mingu Kwak +13

As large language models (LLMs) increasingly generate and process clinical text, scalable evaluation has become critical. LLM-as-a-Judge (LaaJ), which uses LLMs to evaluate model o…

cs.AI2026

Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning

Hang Zhang, Ruheng Wang, Yuelyu Ji +7

Large language models have achieved strong performance on medical reasoning benchmarks, yet their deployment in clinical settings demands rigorous verification to ensure factual ac…

cs.CL2026

MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation

Yuelyu Ji, Min Gu Kwak, Hang Zhang +3

Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with…

cs.AI2025

Orchestrator Multi-Agent Clinical Decision Support System for Secondary Headache Diagnosis in Primary Care

Xizhi Wu, Nelly Estefanie Garduno-Rapp, Justin F Rousseau +6

Unlike most primary headaches, secondary headaches need specialized care and can have devastating consequences if not treated promptly. Clinical guidelines highlight several 'red f…

cs.AI2025

Generative Foundation Model for Structured and Unstructured Electronic Health Records

Sonish Sivarajkumar, Hang Zhang, Yuelyu Ji +6

Electronic health records (EHRs) are rich clinical data sources but complex repositories of patient data, spanning structured elements (demographics, vitals, lab results, codes), u…

cs.CL2025

DeepRAG: Integrating Hierarchical Reasoning and Process Supervision for Biomedical Multi-Hop QA

Yuelyu Ji, Hang Zhang, Shiven Verma +4

We propose DeepRAG, a novel framework that integrates DeepSeek hierarchical question decomposition capabilities with RAG Gym unified retrieval-augmented generation optimization usi…