1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Yusuke Hirota, Ryo Hachiuma, Chao-Han Huck Yang +1
Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual…