7 papers
Self-supervision drives representational convergence in medical foundation models more than clinical supervision
Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia +4
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations on…
Information-seeking failures of large language models in agentic clinical reasoning
Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann +11
Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an…
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan +4
LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-imp…
Benchmarking Foundation Models for Renal Lesion Stratification in CT
Hartmut Häntze, Sarah de Boer, Myrthe Buser +7
The rapid proliferation of open-source medical foundation models (FMs) raises a practical question: how well do their pre-trained representations transfer to clinically relevant bu…
Atomic Fact-Checking Increases Clinician Trust in Large Language Model Recommendations for Oncology Decision Support: A Randomized Controlled Trial
Lisa C. Adams, Linus Marx, Erik Thiele Orberg +8
Question: Does atomic fact-checking, which decomposes AI treatment recommendations into individually verifiable claims linked to source guideline documents, increase clinician trus…
Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering
Mina Farajiamiri, Jeta Sopa, Saba Afza +9
Agentic retrieval-augmented reasoning pipelines are increasingly used to structure how large language models (LLMs) incorporate external evidence in clinical decision support. Thes…