2 citations · 4 across the 8 of their papers we have counts for
4 papers · 1 filter
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
Rezarta Islamaj, Robert Leaman, Joey Chan +13
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabil…
Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
Anran Li, Lingfei Qian, Mengmeng Du +18
Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-training or fine-tuning on medical data to e…
Supervising the search process produces reliable and generalizable information-seeking agents
Guangzhi Xiong, Qiao Jin, Xiao Wang +9
Large language models (LLMs) are transforming web search by shifting from document ranking to synthesizing answers, and are increasingly deployed as autonomous agentic search syste…
AgentMD: Empowering Language Agents for Risk Prediction with Large-Scale Clinical Tool Learning
Qiao Jin, Zhizheng Wang, Yifan Yang +8
Clinical calculators play a vital role in healthcare by offering accurate evidence-based predictions for various purposes such as prognosis. Nevertheless, their widespread utilizat…