11 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution
Qiao Jin, Yin Fang, Lauren He +12
Assessing whether an article supports an assertion is essential for hallucination detection and claim verification. While large language models (LLMs) have the potential to automat…
Entry-level guide to the use of large language models for medical research
Qiao Jin, Nicholas Wan, Robert Leaman +20
Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4, and DeepSeek-R1, represent a transformative class of AI tools capable of revolutionizing variou…
Supervising the search process produces reliable and generalizable information-seeking agents
Guangzhi Xiong, Qiao Jin, Xiao Wang +9
Large language models (LLMs) are transforming web search by shifting from document ranking to synthesizing answers, and are increasingly deployed as autonomous agentic search syste…
DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research
Zifeng Wang, Zheng Chen, Ziwei Yang +5
Biomedical knowledge graphs (KGs) encode vast, heterogeneous information spanning literature, genes, pathways, drugs, diseases, and clinical trials, but leveraging them collectivel…
Developing Large Language Models for Clinical Research Using One Million Clinical Trials
Zifeng Wang, Jiacheng Lin, Qiao Jin +5
Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation that supports model training and rigorous evaluation. Here, we introduce Tria…