13 citations · 13 across the 1 of their papers we have counts for
1 paper
Bang Yang, Fenglin Liu, Xian Wu +3
Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for trainin…