Showing 2024Show all
3 papers · 1 filter
cs.CV2024
A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models
Yuhan Liang, Yijun Li, Yumeng Niu +2
The robustness of Vision-Language Models (VLMs) such as CLIP is critical for their deployment in safety-critical applications like autonomous driving, healthcare diagnostics, and s…
cs.CV2024
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
Yuheng Li, Haotian Liu, Mu Cai +5
In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language mo…
cs.CV2024
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
Zhipeng Bao, Yijun Li, Krishna Kumar Singh +2
Despite recent significant strides achieved by diffusion-based Text-to-Image (T2I) models, current systems are still less capable of ensuring decent compositional generation aligne…