10 papers
Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs
Shayan Mohammadizadehsamakosh, Pritam Sarkar, Leonid Sigal +2
Large Vision-Language Models (LVLMs) have achieved strong performance across medical imaging tasks, yet they remain prone to factual inconsistencies, poor visual grounding, and mis…
Leveraging Foundation Models for Calibration-Free c-VEP BCIs
Mohammadreza Behboodi, Eli Kinney-Lang, Ali Etemad +2
Foundation Models (FMs) have surged in popularity over the past five years, with applications spanning fields from computer vision to natural language processing. Brain-Computer In…
Consistency-Guided Asynchronous Contrastive Tuning for Few-Shot Class-Incremental Tuning of Foundation Models
Shuvendu Roy, Elham Dolatabadi, Arash Afkanpour +1
We propose Consistency-guided Asynchronous Contrastive Tuning (CoACT), a novel method for continuously tuning foundation models to learn new classes in few-shot settings. CoACT con…
Advancing Medical Representation Learning Through High-Quality Data
Negin Baghbanzadeh, Adibvafa Fallahpour, Yasaman Parhizkar +8
Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medi…
A Shared Encoder Approach to Multimodal Representation Learning
Shuvendu Roy, Franklin Ogidi, Ali Etemad +2
Multimodal representation learning has demonstrated remarkable potential in enabling models to process and integrate diverse data modalities, such as text and images, for improved…
Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
Barys Liskavets, Shuvendu Roy, Maxim Ushakov +3
The rise of Large Language Models (LLMs) has led to significant interest in prompt compression, a technique aimed at reducing the length of input prompts while preserving critical…