Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
arXiv:1801.10121
Abstract
Automatic image captioning has recently approached human-level performance due to the latest advances in computer vision and natural language understanding. However, most of the current models can only generate plain factual descriptions about the content of a given image. However, for human beings, image caption writing is quite flexible and diverse, where additional language dimensions, such as emotion, humor and language styles, are often incorporated to produce diverse, emotional, or appealing captions. In particular, we are interested in generating sentiment-conveying image descriptions, which has received little attention. The main challenge is how to effectively inject sentiments into the generated captions without altering the semantic matching between the visual content and the generated descriptions. In this work, we propose two different models, which employ different schemes for injecting sentiments into image captions. Compared with the few existing approaches, the proposed models are much simpler and yet more effective. The experimental results show that our model outperform the state-of-the-art models in generating sentimental (i.e., sentiment-bearing) image captions. In addition, we can also easily manipulate the model by assigning different sentiments to the testing image to generate captions with the corresponding sentiments.
8 pages, 5 figures and 4 tables
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Machine Translation by Jointly Learning to Align and Translate
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Learning to Generate Reviews and Discovering Sentiment
- Image Captioning with Semantic Attention
- Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks
- Toward Controlled Generation of Text
- Attention Correctness in Neural Image Captioning
- Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
- A Hybrid Convolutional Variational Autoencoder for Text Generation
Cited by in corpus (10)
- Controllable Video Captioning with an Exemplar Sentence
- Similar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning
- Senti-Attend: Image Captioning using Sentiment and Attention
- LCEval: Learned Composite Metric for Caption Evaluation
- Stock Price Prediction Under Anomalous Circumstances
- Gaussian Smoothen Semantic Features (GSSF) -- Exploring the Linguistic Aspects of Visual Captioning in Indian Languages (Bengali) Using MSCOCO Framework
- Face-Cap: Image Captioning using Facial Expression Analysis
- Diverse and Styled Image Captioning Using SVD-Based Mixture of Recurrent Experts
- Goal-driven text descriptions for images
- Syntax Customized Video Captioning by Imitating Exemplar Sentences