4 papers
Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
Negin Baghbanzadeh, Mohammed Saidul Islam, Sajad Ashkezari +2
In biomedical vision-language modeling, datasets are typically mined from scientific literature, pairing compound figures with captions that are short, context-dependent, and ofter…
A Flexible Fairness Framework with Surrogate Loss Reweighting for Addressing Sociodemographic Disparities
Wen Xu, Elham Dolatabadi
This paper presents a new algorithmic fairness framework called - Fair Machine Learning (- FML), designed to optimize fairne…
Advancing Medical Representation Learning Through High-Quality Data
Negin Baghbanzadeh, Adibvafa Fallahpour, Yasaman Parhizkar +8
Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medi…
A Shared Encoder Approach to Multimodal Representation Learning
Shuvendu Roy, Franklin Ogidi, Ali Etemad +2
Multimodal representation learning has demonstrated remarkable potential in enabling models to process and integrate diverse data modalities, such as text and images, for improved…