8 papers
RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding
Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu +15
Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, s…
Routine laboratory trajectories encode the onset of organ-level complications in cancer
Jannik Lübberstedt, Krischan Braitsch, Jacqueline Lammert +21
Routine laboratory panels drawn during cancer treatment constitute longitudinal physiological recordings of organ function, yet their temporal structure is discarded by single-time…
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan +4
LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-imp…
Agentic clinical reasoning over longitudinal myeloma records: a retrospective evaluation against expert consensus
Johannes Moll, Jannik Lübberstedt, Christoph Nuernbergk +21
Multiple myeloma is managed through sequential lines of therapy over years to decades, with each decision depending on cumulative disease history distributed across dozens to hundr…
Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve
Weixiang Shen, Bailiang Jian, Jun Li +6
Tool-augmented large language model (LLM) agents can orchestrate specialist classifiers, segmentation models, and visual question-answering modules to interpret chest X-rays. Howev…
Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
Johannes Moll, Markus Graf, Tristan Lemke +7
Vision-language models (VLMs) often produce chain-of-thought (CoT) explanations that sound plausible yet fail to reflect the underlying decision process, undermining trust in high-…