12 papers
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
Jinge Wu, Hongjian Zhou, Mingde Zeng +8
Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness an…
A Regime Theory of Controller Class Selection for LLM Action Decisions
Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim +4
Deployed language and vision-language models must decide, on each input, whether to answer directly, retrieve evidence, defer to a stronger model, or abstain. Contrary to the commo…
RiskAgent: Synergizing Language Models with Validated Tools for Evidence-Based Risk Prediction
Fenglin Liu, Jinge Wu, Hongjian Zhou +9
Large Language Models (LLMs) achieve competitive results compared to human experts in medical examinations. However, it remains a challenge to apply LLMs to complex clinical decisi…
Error Correction in Radiology Reports: A Knowledge Distillation-Based Multi-Stage Framework
Jinge Wu, Zhaolong Wu, Ruizhe Li +6
The increasing complexity and workload of clinical radiology leads to inevitable oversights and mistakes in their use as diagnostic tools, causing delayed treatments and sometimes…
Graph-based LLM over Semi-Structured Population Data for Dynamic Policy Response
Daqian Shi, Xiaolei Diao, Jinge Wu +4
Timely and accurate analysis of population-level data is crucial for effective decision-making during public health emergencies such as the COVID-19 pandemic. However, the massive…
HARE: an entity and relation centric evaluation framework for histopathology reports
Yunsoo Kim, Michal W. S. Ong, Alex Shavick +2
Medical domain automated text generation is an active area of research and development; however, evaluating the clinical quality of generated reports remains a challenge, especiall…