collaborators

5 papers

cs.LG2026

A Comprehensive Evaluation of LLM Reasoning: From Single-Model to Multi-Agent Paradigms

Yapeng Li, Jiakuo Yu, Zhixin Liu +4

Large Language Models (LLMs) are increasingly deployed as reasoning systems, where reasoning paradigms - such as Chain-of-Thought (CoT) and multi-agent systems (MAS) - play a criti…

cs.CL2025

ClinDEF: A Dynamic Evaluation Framework for Large Language Models in Clinical Reasoning

Yuqi Tang, Jing Yu, Zichang Su +7

Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through p…

cs.AI2025

SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration

Keyan Ding, Jing Yu, Junjie Huang +3

Scientific research increasingly relies on specialized computational tools, yet effectively utilizing these tools demands substantial domain expertise. While Large Language Models…

cs.CL2025

OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases

Yongrui Chen, Zhiqiang Liu, Jing Yu +21

Large Language Models (LLMs) have demonstrated substantial progress on reasoning tasks involving unstructured text, yet their capabilities significantly deteriorate when reasoning…

cs.CL2025

SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models

Jing Yu, Yuqi Tang, Kehua Feng +8

Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains r…