2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Junjie Fei, Teng Wang, Jinrui Zhang +3
Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language…