Publications (11)
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
Yixiong Chen, Shawn Xu, Andrew Sellergren +6
Vision-language models have proven to be of great benefit for medical image analysis since they learn rich semantics from both images and reports. Prior efforts have focused on bet…
FastVentricle: Cardiac Segmentation with ENet
Jesse Lieman-Sifry, Matthieu Le, Felix Lau +2
Cardiac Magnetic Resonance (CMR) imaging is commonly used to assess cardiac structure and function. One disadvantage of CMR is that post-processing of exams is tedious. Without aut…
PolyPath: Adapting a Large Multimodal Model for Multi-slide Pathology Report Generation
Faruk Ahmed, Lin Yang, Tiam Jaroensri +11
The interpretation of histopathology cases underlies many important diagnostic and treatment decisions in medicine. Notably, this process typically requires pathologists to integra…
ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders
Shawn Xu, Lin Yang, Christopher Kelly +25
In this work, we present an approach, which we call Embeddings for Language/Image-aligned X-Rays, or ELIXR, that leverages a language-aligned image encoder combined or grafted onto…
ScarGAN: Chained Generative Adversarial Networks to Simulate Pathological Tissue on Cardiovascular MR Scans
Felix Lau, Tom Hendriks, Jesse Lieman-Sifry +3
Medical images with specific pathologies are scarce, but a large amount of data is usually required for a deep convolutional neural network (DCNN) to achieve good accuracy. We cons…
Health AI Developer Foundations
Atilla P. Kiraly, Sebastien Baur, Kenneth Philbrick +23
Robust medical Machine Learning (ML) models have the potential to revolutionize healthcare by accelerating clinical research, improving workflows and outcomes, and producing novel…
Advancing Multimodal Medical Capabilities of Gemini
Lin Yang, Shawn Xu, Andrew Sellergren +44
Many clinical tasks require an understanding of specialized data, such as medical images and genomics, which is not typically found in general-purpose large multimodal models. Buil…
PathAlign: A vision-language model for whole slide images in histopathology
Faruk Ahmed, Andrew Sellergren, Lin Yang +14
Microscopic interpretation of histopathology images underlies many important diagnostic and treatment decisions. While advances in vision-language modeling raise new opportunities…
MedGemma Technical Report
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri +78
Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks,…
MedGemma 1.5 Technical Report
Andrew Sellergren, Chufan Gao, Fereshteh Mahvar +39
We introduce MedGemma 1.5 4B, the latest model in the MedGemma collection. MedGemma 1.5 expands on MedGemma 1 by integrating additional capabilities: high-dimensional medical imagi…
Computationally efficient cardiac views projection using 3D Convolutional Neural Networks
Matthieu Le, Jesse Lieman-Sifry, Felix Lau +3
4D Flow is an MRI sequence which allows acquisition of 3D images of the heart. The data is typically acquired volumetrically, so it must be reformatted to generate cardiac long axi…