1 citations · 1 across the 5 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation
Fuyi Yang, Chenchen Ye, Mingyu Derek Ma +3
Hypothesis generation in biomedical research has traditionally centered on uncovering hidden relationships within vast scientific literature, often using methods like Literature-Ba…
cs.CL2024★ 1 cited
MIRAI: Evaluating LLM Agents for Event Forecasting
Chenchen Ye, Ziniu Hu, Yihe Deng +4
Recent advancements in Large Language Models (LLMs) have empowered LLM agents to autonomously collect world information, over which to conduct reasoning to solve complex problems.…
cs.CL2024
CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making
Mingyu Derek Ma, Chenchen Ye, Yu Yan +4
The integration of Artificial Intelligence (AI), especially Large Language Models (LLMs), into the clinical diagnosis process offers significant potential to improve the efficiency…