6 papers
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
Ajmal M., Abin Roy, Afthab Salam Kanniyan +4
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…
Developing Large Language Models for Clinical Research Using One Million Clinical Trials
Zifeng Wang, Jiacheng Lin, Qiao Jin +5
Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation that supports model training and rigorous evaluation. Here, we introduce Tria…
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
Zifeng Wang, Benjamin Danek, Jimeng Sun
Validating scientific hypotheses is a central challenge in biomedical research, and remains difficult for artificial intelligence (AI) agents due to the complexity of real-world da…
Can Large Language Models Replace Data Scientists in Biomedical Research?
Zifeng Wang, Benjamin Danek, Ziwei Yang +2
Data science plays a critical role in biomedical research, but it requires professionals with expertise in coding and medical data analysis. Large language models (LLMs) have shown…
InformGen: An AI Copilot for Accurate and Compliant Clinical Research Consent Document Generation
Zifeng Wang, Junyi Gao, Benjamin Danek +5
Leveraging large language models (LLMs) to generate high-stakes documents, such as informed consent forms (ICFs), remains a significant challenge due to the extreme need for regula…
A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges
Zifeng Wang, Hanyin Wang, Benjamin Danek +6
The integration of Large Language Models (LLMs) into medical applications has sparked widespread interest across the healthcare industry, from drug discovery and development to cli…