collaborators

6 papers

cs.AI2026

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

Jinge Wu, Hongjian Zhou, Mingde Zeng +8

Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness an…

cs.CL2026

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Hongjian Zhou, Xinyu Zou, Jinge Wu +19

Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…

cs.AI2026

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI

Sean Wu, Pan Lu, Yupeng Chen +7

AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances…

cs.CL2026

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence

Sean Wu, Fredrik K. Gustafsson, Edward Phillips +3

Large language models (LLMs) often produce confident but incorrect answers in settings where abstention would be safer. Standard evaluation protocols, however, require a response a…

cs.CL2026

Entropy Alone is Insufficient for Safe Selective Prediction in LLMs

Edward Phillips, Fredrik K. Gustafsson, Sean Wu +2

Selective prediction systems can mitigate harms resulting from language model hallucinations by abstaining from answering in high-risk cases. Uncertainty quantification techniques…

cs.CL2025

AutoMedPrompt: A New Framework for Optimizing LLM Medical Prompts Using Textual Gradients

Sean Wu, Michael Koo, Fabien Scalzo +1

Large language models (LLMs) have demonstrated increasingly sophisticated performance in medical and other fields of knowledge. Traditional methods of creating specialist LLMs requ…