32 citations · 33 across the 3 of their papers we have counts for
5 papers · 1 filter
Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs
Shayan Mohammadizadehsamakosh, Pritam Sarkar, Leonid Sigal +2
Large Vision-Language Models (LVLMs) have achieved strong performance across medical imaging tasks, yet they remain prone to factual inconsistencies, poor visual grounding, and mis…
A Shared Encoder Approach to Multimodal Representation Learning
Shuvendu Roy, Franklin Ogidi, Ali Etemad +2
Multimodal representation learning has demonstrated remarkable potential in enabling models to process and integrate diverse data modalities, such as text and images, for improved…
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
Shuvendu Roy, Ali Etemad
We present SelfPrompt, a novel prompt-tuning approach for vision-language models (VLMs) in a semi-supervised learning setup. Existing methods for tuning VLMs in semi-supervised set…
Self-supervised Contrastive Learning of Multi-view Facial Expressions
Shuvendu Roy, Ali Etemad
Facial expression recognition (FER) has emerged as an important component of human-computer interaction systems. Despite recent advancements in FER, performance often drops signifi…
Spatiotemporal Contrastive Learning of Facial Expressions in Videos
Shuvendu Roy, Ali Etemad
We propose a self-supervised contrastive learning approach for facial expression recognition (FER) in videos. We propose a novel temporal sampling-based augmentation scheme to be u…