7 papers
Accelerating Reproducible Research in Synthetic EHR Generation
Jalen Jiang, Chufan Gao, Ethan Rasmussen +2
The generation of high-fidelity synthetic Electronic Health Records (EHR) is crucial for advancing medical research while preserving patient privacy. However, head-to-head comparis…
MedGemma 1.5 Technical Report
Andrew Sellergren, Chufan Gao, Fereshteh Mahvar +39
We introduce MedGemma 1.5 4B, the latest model in the MedGemma collection. MedGemma 1.5 expands on MedGemma 1 by integrating additional capabilities: high-dimensional medical imagi…
Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise
Hanyin Wang, Chufan Gao, Qiping Xu +9
Process-supervised reward models (PRMs) excel at providing step-by-step verification for large language model (LLM) outputs in domains like mathematics and coding. However, their a…
Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding
Chufan Gao, Jintai Chen, Jimeng Sun
Automated tabular understanding and reasoning are essential tasks for data scientists. Recently, Large language models (LLMs) have become increasingly prevalent in tabular reasonin…
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
Hanyin Wang, Chufan Gao, Bolun Liu +7
Proprietary Large Language Models (LLMs) such as GPT-4 and Gemini have demonstrated promising capabilities in clinical text summarization tasks. However, due to patient data privac…
Automatically Labeling Clinical Trial Outcomes: A Large-Scale Benchmark for Drug Development
Chufan Gao, Jathurshan Pradeepkumar, Trisha Das +2
Background The cost of drug discovery and development is substantial, with clinical trial outcomes playing a critical role in regulatory approval and patient care. However, access…