2 papers
cs.CV2026
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
Ashwin Kumar, Robbie Holland, Corey Barrett +8
Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-style finetuning. This two-sta…
cs.CV2026
Sparse Autoencoders for Interpretable Medical Image Representation Learning
Philipp Wesp, Robbie Holland, Vasiliki Sideri-Lampretsa +1
Vision foundation models (FMs) achieve state-of-the-art performance in medical imaging. However, they encode information in abstract latent representations that clinicians cannot i…