5 papers
A Comprehensive Evaluation of LLM Reasoning: From Single-Model to Multi-Agent Paradigms
Yapeng Li, Jiakuo Yu, Zhixin Liu +4
Large Language Models (LLMs) are increasingly deployed as reasoning systems, where reasoning paradigms - such as Chain-of-Thought (CoT) and multi-agent systems (MAS) - play a criti…
ClinDEF: A Dynamic Evaluation Framework for Large Language Models in Clinical Reasoning
Yuqi Tang, Jing Yu, Zichang Su +7
Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through p…
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
Keyan Ding, Jing Yu, Junjie Huang +3
Scientific research increasingly relies on specialized computational tools, yet effectively utilizing these tools demands substantial domain expertise. While Large Language Models…
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
Yongrui Chen, Zhiqiang Liu, Jing Yu +21
Large Language Models (LLMs) have demonstrated substantial progress on reasoning tasks involving unstructured text, yet their capabilities significantly deteriorate when reasoning…
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
Jing Yu, Yuqi Tang, Kehua Feng +8
Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains r…