activity
20242026
collaborators

6 papers

cs.CL2026

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Ajmal M., Abin Roy, Afthab Salam Kanniyan +4

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…

cs.AI2025

Developing Large Language Models for Clinical Research Using One Million Clinical Trials

Zifeng Wang, Jiacheng Lin, Qiao Jin +5

Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation that supports model training and rigorous evaluation. Here, we introduce Tria…

cs.AI2025

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

Zifeng Wang, Benjamin Danek, Jimeng Sun

Validating scientific hypotheses is a central challenge in biomedical research, and remains difficult for artificial intelligence (AI) agents due to the complexity of real-world da…

cs.AI2025

Can Large Language Models Replace Data Scientists in Biomedical Research?

Zifeng Wang, Benjamin Danek, Ziwei Yang +2

Data science plays a critical role in biomedical research, but it requires professionals with expertise in coding and medical data analysis. Large language models (LLMs) have shown…

cs.CL2025

InformGen: An AI Copilot for Accurate and Compliant Clinical Research Consent Document Generation

Zifeng Wang, Junyi Gao, Benjamin Danek +5

Leveraging large language models (LLMs) to generate high-stakes documents, such as informed consent forms (ICFs), remains a significant challenge due to the extreme need for regula…

cs.CL2024

A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges

Zifeng Wang, Hanyin Wang, Benjamin Danek +6

The integration of Large Language Models (LLMs) into medical applications has sparked widespread interest across the healthcare industry, from drug discovery and development to cli…