SuperCaptioning: Image Captioning Using Two-dimensional Word Embedding
arXiv:1905.10515
Abstract
Language and vision are processed as two different modal in current work for image captioning. However, recent work on Super Characters method shows the effectiveness of two-dimensional word embedding, which converts text classification problem into image classification problem. In this paper, we propose the SuperCaptioning method, which borrows the idea of two-dimensional word embedding from Super Characters method, and processes the information of language and vision together in one single CNN model. The experimental results on Flickr30k data shows the proposed method gives high quality image captions. An interactive demo is ready to show at the workshop.
3 pages, 2 figures, modified typo. Accepted by CVPR2019 VQA workshop
References in corpus (4)
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Squared English Word: A Method of Generating Glyph to Use Super Characters for Sentiment Analysis
- SuperTML: Two-Dimensional Word Embedding for the Precognition on Structured Tabular Data
- SuperChat: Dialogue Generation by Transfer Learning from Vision to Language using Two-dimensional Word Embedding and Pretrained ImageNet CNN Models