4 papers
Fine-grained CLIP fine-tuning with self-annotated region alignment
Chenyang Zhao, Wei Lin, Janet H. Hsiao +1
Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the…
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
Xin Huang, Antoni B. Chan
Large Language Models (LLMs) are increasingly evaluated with input attribution methods, yet comparing such explanations remains challenging. Existing soft-perturbation faithfulness…
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
Jiuniu Wang, Wenjia Xu, Qingzhong Wang +1
Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high per…
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
Chenyang Zhao, Kun Wang, Janet H. Hsiao +1
Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is…