12 papers · 1 filter
ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
Canyu Chen, Jian Yu, Shan Chen +8
Large Language Models (LLMs) hold great promise to revolutionize current clinical systems for their superior capacities on medical text processing tasks and medical licensing exams…
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
Natalie Seah, Danielle S. Bitterman, Daphna Spiegel +1
Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where clinicians must clearly and…
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Bingyang Ye, Shan Chen, Jingxuan Tu +4
Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientif…
Simulated patient systems powered by large language model-based AI agents offer potential for transforming medical education
Huizi Yu, Jiayan Zhou, Lingyao Li +22
Background: Simulated patient systems are important in medical education and research, providing safe, integrative training environments and supporting clinical decision making. Ad…
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
Jirui Qi, Shan Chen, Zidi Xiong +3
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studi…
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
Shan Chen, Pedro Moreira, Yuxin Xiao +6
Large language models (LLMs) are increasingly envisioned as decision-support tools in clinical practice, yet safe clinical reasoning demands integrating heterogeneous knowledge bas…