3 papers
cs.MM2024
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
Mianzhi Pan, Jianfei Li, Mingyue Yu +4
Commonsense reasoning, the ability to make logical assumptions about daily scenes, is one core intelligence of human beings. In this work, we present a novel task and dataset for e…
cs.CV2023
ADS-Cap: A Framework for Accurate and Diverse Stylized Captioning with Unpaired Stylistic Corpora
Kanzhi Cheng, Zheng Ma, Shi Zong +3
Generating visually grounded image captions with specific linguistic styles using unpaired stylistic corpora is a challenging task, especially since we expect stylized captions wit…
cs.CV2023
Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training Model
Kanzhi Cheng, Wenpo Song, Zheng Ma +3
Current captioning approaches tend to generate correct but "generic" descriptions that lack real-world knowledge, e.g., named entities and contextual information. Considering that…