Multimodal Word Distributions
arXiv:1704.08424
Abstract
Word embeddings provide point representations of words containing useful semantic information. We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information. To learn these distributions, we propose an energy-based max-margin objective. We show that the resulting approach captures uniquely expressive semantic information, and outperforms alternatives, such as word2vec skip-grams, and Gaussian embeddings, on benchmark datasets such as word similarity and entailment.
This paper also appears at ACL 2017
References in corpus (3)
Cited by in corpus (11)
- Generative Models of Visually Grounded Imagination
- Generalizing Point Embeddings using the Wasserstein Space of Elliptical Distributions
- Embedding Words as Distributions with a Bayesian Skip-gram Model
- Robust Multilingual Part-of-Speech Tagging via Adversarial Training
- Gaussian Word Embedding with a Wasserstein Distance Loss
- Mixture-of-tastes Models for Representing Users with Diverse Interests
- Efficient Graph-based Word Sense Induction by Distributional Inclusion Vector Embeddings
- Learning Taxonomies of Concepts and not Words using Contextualized Word Representations: A Position Paper
- Using Multi-Sense Vector Embeddings for Reverse Dictionaries
- Exploration on Grounded Word Embedding: Matching Words and Images with Image-Enhanced Skip-Gram Model
- Learning Multi-Sense Word Distributions using Approximate Kullback-Leibler Divergence