Order-Embeddings of Images and Language
arXiv:1511.06361
Abstract
Hypernymy, textual entailment, and image captioning can be seen as special cases of a single visual-semantic hierarchy over words, sentences, and images. In this paper we advocate for explicitly modeling the partial order structure of this hierarchy. Towards this goal, we introduce a general method for learning ordered representations, and show how it can be applied to a variety of tasks involving images and language. We show that the resulting representations improve performance over current approaches for hypernym prediction and image-caption retrieval.
ICLR camera-ready version
References in corpus (3)
Cited by in corpus (43)
- Enhanced LSTM for Natural Language Inference
- VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
- An efficient framework for learning sentence representations
- X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data
- Learning Natural Language Inference using Bidirectional LSTM model and Inner-Attention
- Poincaré Embeddings for Learning Hierarchical Representations
- Representations of language in a model of visually grounded speech signal
- Bilateral Multi-Perspective Matching for Natural Language Sentences
- Root Mean Square Layer Normalization
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- Distance-based Self-Attention Network for Natural Language Inference
- Learning Natural Language Inference with LSTM
- A Decomposable Attention Model for Natural Language Inference
- Emergent Translation in Multi-Agent Communication
- Natural Language Inference by Tree-Based Convolution and Heuristic Matching
- Hyperbolic Entailment Cones for Learning Hierarchical Embeddings
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval
- Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning
- Knowledge Graph Embedding with Iterative Guidance from Soft Rules
- On Learning Associations of Faces and Voices
- Fishing for Clickbaits in Social Images and Texts with Linguistically-Infused Neural Network Models
- Large-Scale Visual Relationship Understanding
- Open-World Visual Recognition Using Knowledge Graphs
- Visual-Textual Association with Hardest and Semi-Hard Negative Pairs Mining for Person Search
- An Efficient Approach to Informative Feature Extraction from Multimodal Data
- A Neural Architecture Mimicking Humans End-to-End for Natural Language Inference
- Scene Graph Parsing by Attention Graph
- Scene Graph Parsing as Dependency Parsing
- Better Text Understanding Through Image-To-Text Transfer
- Path-Based Contextualization of Knowledge Graphs for Textual Entailment
- What If We Simply Swap the Two Text Fragments? A Straightforward yet Effective Way to Test the Robustness of Methods to Confounding Signals in Nature Language Inference Tasks
- Disjoint Multi-task Learning between Heterogeneous Human-centric Tasks
- Not All Words are Equal: Video-specific Information Loss for Video Captioning
- Lessons learned in multilingual grounded language learning
- Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks
- Understanding Roles and Entities: Datasets and Models for Natural Language Inference
- Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations
- Contextually Customized Video Summaries via Natural Language
- Double Forward Propagation for Memorized Batch Normalization
- YouMakeup VQA Challenge: Towards Fine-grained Action Understanding in Domain-Specific Videos
- Embedding Geographic Locations for Modelling the Natural Environment using Flickr Tags and Structured Data
- Show, Translate and Tell