From the 1 of 4 linked papers with an AI index.
4 papers
Fine-grained CLIP fine-tuning with self-annotated region alignment
Chenyang Zhao, Wei Lin, Antoni B. Chan +1
The paper proposes SFF-CLIP, a fine-tuning approach that uses only image-text pairs to align region features with phrase concepts via text-specific heat maps, improving CLIP's fine…
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
Xin Huang, Antoni B. Chan
Large Language Models (LLMs) are increasingly evaluated with input attribution methods, yet comparing such explanations remains challenging. Existing soft-perturbation faithfulness…
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
Chenyang Zhao, Kun Wang, Janet H. Hsiao +1
Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is…
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
Jiuniu Wang, Wenjia Xu, Qingzhong Wang +1
Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high per…