2 papers
cs.AI2026
Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy
Yuan Shen, Xiaojun Wu, Linghua Yu
Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we present AI-MASLD, a stress-au…
cs.AI2025
AI-MASLD Metabolic Dysfunction and Information Steatosis of Large Language Models in Unstructured Clinical Narratives
Yuan Shen, Xiaojun Wu, Linghua Yu
This study aims to simulate real-world clinical scenarios to systematically evaluate the ability of Large Language Models (LLMs) to extract core medical information from patient ch…