Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
Yujin Jo, Sangyoon Bae, Taesup Kim
Hallucinations in large vision--language models (LVLMs) often arise when language priors dominate over visual evidence, leading to object misidentification and visually inconsisten…
cs.CV2024
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling
Donggeun Kim, Yujin Jo, Myungjoo Lee +1
The advancement of vision-language models, particularly the Contrastive Language-Image Pre-training (CLIP) model, has revolutionized the field of machine learning by enabling robus…
cs.CV2024
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
Donggeun Kim, Taesup Kim
Multimodal learning typically relies on the assumption that all modalities are fully available during both the training and inference phases. However, in real-world scenarios, cons…