3 papers
cs.CV2025
Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
Negin Baghbanzadeh, Mohammed Saidul Islam, Sajad Ashkezari +2
In biomedical vision-language modeling, datasets are typically mined from scientific literature, pairing compound figures with captions that are short, context-dependent, and ofter…
eess.IV2025
Advancing Medical Representation Learning Through High-Quality Data
Negin Baghbanzadeh, Adibvafa Fallahpour, Yasaman Parhizkar +8
Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medi…
cs.CV2025
Similarity-Aware Token Pruning: Your VLM but Faster
Ahmadreza Jeddi, Negin Baghbanzadeh, Elham Dolatabadi +1
The computational demands of Vision Transformers (ViTs) and Vision-Language Models (VLMs) remain a significant challenge due to the quadratic complexity of self-attention. While to…