7 papers
On the Diversity of Analogy Making in Large Language Models
Yuanhao Shen, Daniel Xavier de Sousa, Caio César Sifuentes Barcelos +3
Large Language Models (LLMs) have demonstrated remarkable potential for analogy making, a core cognitive capability that drives novelty and creativity. While prior research has ext…
Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy
Yuan Shen, Xiaojun Wu, Linghua Yu
Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we present AI-MASLD, a stress-au…
IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research
Yuanhao Shen, Daniel Xavier de Sousa, Ricardo Marçal +2
Innovation is a key driving force of human civilization. As the body of knowledge has grown considerably, bridging knowledge across different disciplines, where significant innovat…
AI-MASLD Metabolic Dysfunction and Information Steatosis of Large Language Models in Unstructured Clinical Narratives
Yuan Shen, Xiaojun Wu, Linghua Yu
This study aims to simulate real-world clinical scenarios to systematically evaluate the ability of Large Language Models (LLMs) to extract core medical information from patient ch…
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
Yufeng Zhao, Junnan Liu, Hongwei Liu +4
Large Language Models (LLMs) have made significant strides in reasoning tasks through methods like chain-of-thought (CoT) reasoning. However, they often fall short in tasks requiri…
Validating the Effectiveness of a Large Language Model-based Approach for Identifying Children's Development across Various Free Play Settings in Kindergarten
Yuanyuan Yang, Yuan Shen, Tianchen Sun +1
Free play is a fundamental aspect of early childhood education, supporting children's cognitive, social, emotional, and motor development. However, assessing children's development…