2 papers
cs.CV2023
Bounding and Filling: A Fast and Flexible Framework for Image Captioning
Zheng Ma, Changxin Wang, Bo Huang +2
Most image captioning models following an autoregressive manner suffer from significant inference latency. Several models adopted a non-autoregressive manner to speed up the proces…
cs.CV2023
Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training Model
Kanzhi Cheng, Wenpo Song, Zheng Ma +3
Current captioning approaches tend to generate correct but "generic" descriptions that lack real-world knowledge, e.g., named entities and contextual information. Considering that…