4 papers
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
Xiao Zhang, Qumeng Sun, Jiahao Li +4
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy st…
AutoResearch: Insight In, Hallucination Out
Yiming Ren, Xiang Liu, Qumeng Sun +4
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically gr…
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
Yuxuan Zhou, Xien Liu, Chenwei Yan +8
Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored.…
Reliable and diverse evaluation of LLM medical knowledge mastery
Yuxuan Zhou, Xien Liu, Chen Ning +2
Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing…