4 papers
ClinDEF: A Dynamic Evaluation Framework for Large Language Models in Clinical Reasoning
Yuqi Tang, Jing Yu, Zichang Su +7
Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through p…
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
Keyan Ding, Jing Yu, Junjie Huang +3
Scientific research increasingly relies on specialized computational tools, yet effectively utilizing these tools demands substantial domain expertise. While Large Language Models…
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
Yongrui Chen, Zhiqiang Liu, Jing Yu +21
Large Language Models (LLMs) have demonstrated substantial progress on reasoning tasks involving unstructured text, yet their capabilities significantly deteriorate when reasoning…
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
Jing Yu, Yuqi Tang, Kehua Feng +8
Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains r…