Improving Visually Grounded Sentence Representations with Self-Attention
arXiv:1712.00609
Abstract
Sentence representation models trained only on language could potentially suffer from the grounding problem. Recent work has shown promising results in improving the qualities of sentence representations by jointly training them with associated image features. However, the grounding capability is limited due to distant connection between input sentences and image features by the design of the architecture. In order to further close the gap, we propose applying self-attention mechanism to the sentence encoder to deepen the grounding effect. Our results on transfer tasks show that self-attentive encoders are better for visual grounding, as they exploit specific words with strong visual associations.
References in corpus (8)
- A Structured Self-attentive Sentence Embedding
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
- Efficient Vector Representation for Documents through Corruption
- Deconvolutional Paragraph Representation Learning
- Learning language through pictures
- Improving Pairwise Ranking for Multi-label Image Classification
- Rethinking Skip-thought: A Neighborhood based Approach