Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
Chenyang Zhao, Kun Wang, Janet H. Hsiao +1
Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is…
cs.CV2025
V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
Nan Sun, Zhenyu Zhang, Xixun Lin +8
Multimodal Large Language Models (MLLMs) excel in numerous vision-language tasks yet suffer from hallucinations, producing content inconsistent with input visuals, that undermine r…