1 paper
Binbin Li, Guimiao Yang, Zisen Qi +2
Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual feat…