3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.CV2021
Boosting Entity-aware Image Captioning with Multi-modal Knowledge Graph
Wentian Zhao, Yao Hu, Heda Wang +2
Entity-aware image captioning aims to describe named entities and events related to the image by utilizing the background knowledge in the associated article. This task remains cha…
cs.MM2020★ 3 cited
LAMP: Label Augmented Multimodal Pretraining
Jia Guo, Chen Zhu, Yilun Zhao +4
Multi-modal representation learning by pretraining has become an increasing interest due to its easy-to-use and potential benefit for various Visual-and-Language~(V-L) tasks. Howev…