Enriching Word Vectors with Subword Information
arXiv:1607.04606
Abstract
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. This is a limitation, especially for languages with large vocabularies and many rare words. In this paper, we propose a new approach based on the skipgram model, where each word is represented as a bag of character -grams. A vector representation is associated to each character -gram; words being represented as the sum of these representations. Our method is fast, allowing to train models on large corpora quickly and allows us to compute word representations for words that did not appear in the training data. We evaluate our word representations on nine different languages, both on word similarity and analogy tasks. By comparing to recently proposed morphological word representations, we show that our vectors achieve state-of-the-art performance on these tasks.
Accepted to TACL. The two first authors contributed equally
References in corpus (2)
Cited by in corpus (202)
- Recent Trends in Deep Learning Based Natural Language Processing
- Enhancing Clinical Concept Extraction with Contextual Embeddings
- Optimal Hyperparameters for Deep LSTM-Networks for Sequence Labeling Tasks
- Natural Language Processing Advancements By Deep Learning: A Survey
- Poincaré Embeddings for Learning Hierarchical Representations
- Checking Smart Contracts with Structural Code Embedding
- Hierarchical Neural Story Generation
- Unsupervised Question Answering by Cloze Translation
- Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- Portuguese Word Embeddings: Evaluating on Word Analogies and Natural Language Tasks
- Unsupervised Hyperalignment for Multilingual Word Embeddings
- Analogical Reasoning on Chinese Morphological and Semantic Relations
- Few-shot Learning for Named Entity Recognition in Medical Text
- Modeling Noisiness to Recognize Named Entities using Multitask Neural Networks on Social Media
- Zero-Shot Action Recognition in Videos: A Survey
- Double Embeddings and CNN-based Sequence Labeling for Aspect Extraction
- Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors
- Joint Entity Extraction and Assertion Detection for Clinical Text
- An Introductory Survey on Attention Mechanisms in NLP Problems
- CASCADE: Contextual Sarcasm Detection in Online Discussion Forums
- Unsupervised Alignment of Embeddings with Wasserstein Procrustes
- SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval)
- What's in a Name? Reducing Bias in Bios without Access to Protected Attributes
- Hex2vec -- Context-Aware Embedding H3 Hexagons with OpenStreetMap Tags
- IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
- Hate Speech Detection on Vietnamese Social Media Text using the Bi-GRU-LSTM-CNN Model
- Amobee at SemEval-2018 Task 1: GRU Neural Network with a CNN Attention Mechanism for Sentiment Classification
- Meta-Learning for Low-Resource Neural Machine Translation
- Enhance word representation for out-of-vocabulary on Ubuntu dialogue corpus
- Beheshti-NER: Persian Named Entity Recognition Using BERT
- A Fine-Grained Sentiment Dataset for Norwegian
- Forecasting the presence and intensity of hostility on Instagram using linguistic and social features
- Extreme Multi-Label Legal Text Classification: A case study in EU Legislation
- Flexible and Scalable State Tracking Framework for Goal-Oriented Dialogue Systems
- Exploration of Neural Machine Translation in Autoformalization of Mathematics in Mizar
- Reuse and Adaptation for Entity Resolution through Transfer Learning
- Named Entity Recognition as Dependency Parsing
- Is preprocessing of text really worth your time for online comment classification?
- Robust Lexical Features for Improved Neural Network Named-Entity Recognition
- Investigating the Effects of Word Substitution Errors on Sentence Embeddings
- Surfacing contextual hate speech words within social media
- Towards Sub-Word Level Compositions for Sentiment Analysis of Hindi-English Code Mixed Text
- Non-Adversarial Unsupervised Word Translation
- Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation Evaluators
- Membership Inference on Word Embedding and Beyond
- Triple Classification for Scholarly Knowledge Graph Completion
- Cross-Dataset Design Discussion Mining
- Evaluating Word Embeddings for Sentence Boundary Detection in Speech Transcripts
- MOLIERE: Automatic Biomedical Hypothesis Generation System
- Clinical Relation Extraction Using Transformer-based Models
- Alquist 2.0: Alexa Prize Socialbot Based on Sub-Dialogue Models
- Sentence Boundary Detection for French with Subword-Level Information Vectors and Convolutional Neural Networks
- Leveraging Subword Embeddings for Multinational Address Parsing
- Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages
- Evolutionary Neural AutoML for Deep Learning
- BHAAV- A Text Corpus for Emotion Analysis from Hindi Stories
- UKP-Athene: Multi-Sentence Textual Entailment for Claim Verification
- Learning Emoji Embeddings using Emoji Co-occurrence Network Graph
- Application of a Hybrid Bi-LSTM-CRF model to the task of Russian Named Entity Recognition
- TableQA: Question Answering on Tabular Data
- Unsupervised Sentiment Analysis for Code-mixed Data
- AttnConvnet at SemEval-2018 Task 1: Attention-based Convolutional Neural Networks for Multi-label Emotion Classification
- Where are the Keys? -- Learning Object-Centric Navigation Policies on Semantic Maps with Graph Convolutional Networks
- Natural language understanding for task oriented dialog in the biomedical domain in a low resources context
- Alquist 3.0: Alexa Prize Bot Using Conversational Knowledge Graph
- Multiscale sequence modeling with a learned dictionary
- Convolutional Neural Networks for Sentiment Classification on Business Reviews
- Review Helpfulness Prediction with Embedding-Gated CNN
- Automated Spelling Correction for Clinical Text Mining in Russian
- Relaxed Softmax for learning from Positive and Unlabeled data
- A Survey of Neural Network Techniques for Feature Extraction from Text
- Cross-Lingual Transfer Learning for Question Answering
- Sentiment Analysis on Financial News Headlines using Training Dataset Augmentation
- Improving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations
- Diving Deep into Clickbaits: Who Use Them to What Extents in Which Topics with What Effects?
- Neutralizing Gender Bias in Word Embedding with Latent Disentanglement and Counterfactual Generation
- Towards Debiasing Sentence Representations
- Sentiment Analysis of Code-Mixed Social Media Text (Hinglish)
- Fast Amortized Inference and Learning in Log-linear Models with Randomly Perturbed Nearest Neighbor Search
- LSTM Easy-first Dependency Parsing with Pre-trained Word Embeddings and Character-level Word Embeddings in Vietnamese
- On the Impact of Knowledge-based Linguistic Annotations in the Quality of Scientific Embeddings
- Efficacy of BERT embeddings on predicting disaster from Twitter data
- JUNLP@Dravidian-CodeMix-FIRE2020: Sentiment Classification of Code-Mixed Tweets using Bi-Directional RNN and Language Tags
- Investigating an Effective Character-level Embedding in Korean Sentence Classification
- Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models
- Always Lurking: Understanding and Mitigating Bias in Online Human Trafficking Detection
- An Exploration of Data Augmentation Techniques for Improving English to Tigrinya Translation
- UFTR: A Unified Framework for Ticket Routing
- Reinforced Data Sampling for Model Diversification
- Query expansion with artificially generated texts
- The Best of Both Worlds: Lexical Resources To Improve Low-Resource Part-of-Speech Tagging
- Advancing PICO Element Detection in Biomedical Text via Deep Neural Networks
- Event detection in Colombian security Twitter news using fine-grained latent topic analysis
- Towards Computational Linguistics in Minangkabau Language: Studies on Sentiment Analysis and Machine Translation
- An Improved Neural Baseline for Temporal Relation Extraction
- The ARIEL-CMU Systems for LoReHLT18
- Attention-based Mixture Density Recurrent Networks for History-based Recommendation
- A Novel Distributed Representation of News (DRNews) for Stock Market Predictions
- Distant Supervision and Noisy Label Learning for Low Resource Named Entity Recognition: A Study on Hausa and Yorùbá
- Topological Data Analysis in Text Classification: Extracting Features with Additive Information
- Seeing The Whole Patient: Using Multi-Label Medical Text Classification Techniques to Enhance Predictions of Medical Codes
- Generalised Differential Privacy for Text Document Processing
- Predicting user intent from search queries using both CNNs and RNNs
- An Analysis of Approaches Taken in the ACM RecSys Challenge 2018 for Automatic Music Playlist Continuation
- Combining Context-Free and Contextualized Representations for Arabic Sarcasm Detection and Sentiment Identification
- Language Modeling by Clustering with Word Embeddings for Text Readability Assessment
- Multi-Label Sentiment Analysis on 100 Languages with Dynamic Weighting for Label Imbalance
- Aff2Vec: Affect--Enriched Distributional Word Representations
- A Novel Method of Extracting Topological Features from Word Embeddings
- Style Obfuscation by Invariance
- Sentiment Classification with Word Attention based on Weakly Supervised Learning with a Convolutional Neural Network
- Combination of multiple Deep Learning architectures for Offensive Language Detection in Tweets
- Important Attribute Identification in Knowledge Graph
- Microsoft AI Challenge India 2018: Learning to Rank Passages for Web Question Answering with Deep Attention Networks
- A Novel Deep Learning Method for Textual Sentiment Analysis
- On Extending NLP Techniques from the Categorical to the Latent Space: KL Divergence, Zipf's Law, and Similarity Search
- Error Analysis for Vietnamese Named Entity Recognition on Deep Neural Network Models
- Expansional Retrofitting for Word Vector Enrichment
- Dependency Grammar Induction with a Neural Variational Transition-based Parser
- Word2Vec is a special case of Kernel Correspondence Analysis and Kernels for Natural Language Processing
- Differentiable Greedy Networks
- Extrofitting: Enriching Word Representation and its Vector Space with Semantic Lexicons
- Discriminating Between Similar Nordic Languages
- The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability
- Learning How to Self-Learn: Enhancing Self-Training Using Neural Reinforcement Learning
- Toward Cross-Lingual Definition Generation for Language Learners
- ASBERT: Siamese and Triplet network embedding for open question answering
- Gaussian Hierarchical Latent Dirichlet Allocation: Bringing Polysemy Back
- Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages
- Entity Candidate Network for Whole-Aware Named Entity Recognition
- Worse WER, but Better BLEU? Leveraging Word Embedding as Intermediate in Multitask End-to-End Speech Translation
- Text Matters but Speech Influences: A Computational Analysis of Syntactic Ambiguity Resolution
- Synapse at CAp 2017 NER challenge: Fasttext CRF
- Question Relevance in Visual Question Answering
- Exploring Semi-supervised Variational Autoencoders for Biomedical Relation Extraction
- Predicting the Argumenthood of English Prepositional Phrases
- Neural Coreference Resolution for Arabic
- Tackling Morphological Analogies Using Deep Learning -- Extended Version
- Multichannel LSTM-CNN for Telugu Technical Domain Identification
- Deep Learning Models in Detection of Dietary Supplement Adverse Event Signals from Twitter
- More Romanian word embeddings from the RETEROM project
- Nonsymbolic Text Representation
- Modeling Institutional Credit Risk with Financial News
- SLAM-Inspired Simultaneous Contextualization and Interpreting for Incremental Conversation Sentences
- Combining word embeddings and convolutional neural networks to detect duplicated questions
- Creation and Evaluation of Datasets for Distributional Semantics Tasks in the Digital Humanities Domain
- Relation Extraction Datasets in the Digital Humanities Domain and their Evaluation with Word Embeddings
- Attending Form and Context to Generate Specialized Out-of-VocabularyWords Representations
- Generating Sense Embeddings for Syntactic and Semantic Analogy for Portuguese
- Italian Event Detection Goes Deep Learning
- IndoSum: A New Benchmark Dataset for Indonesian Text Summarization
- Sketching Transformed Matrices with Applications to Natural Language Processing
- Hierarchical Neural Networks for Sequential Sentence Classification in Medical Scientific Abstracts
- Translations as Additional Contexts for Sentence Classification
- UMDSub at SemEval-2018 Task 2: Multilingual Emoji Prediction Multi-channel Convolutional Neural Network on Subword Embedding
- VCWE: Visual Character-Enhanced Word Embeddings
- Word Embeddings for the Armenian Language: Intrinsic and Extrinsic Evaluation
- Named Entity Recognition on Code-Switched Data: Overview of the CALCS 2018 Shared Task
- ALL-IN-1: Short Text Classification with One Model for All Languages
- The Interplay of Semantics and Morphology in Word Embeddings
- ThamizhiUDp: A Dependency Parser for Tamil
- ODSQA: Open-domain Spoken Question Answering Dataset
- Automated Extraction of Personal Knowledge from Smartphone Push Notifications
- Efficient Purely Convolutional Text Encoding
- Morphological Skip-Gram: Using morphological knowledge to improve word representation
- Applying recent advances in Visual Question Answering to Record Linkage
- Interpretable Structure-aware Document Encoders with Hierarchical Attention
- Phonetic-and-Semantic Embedding of Spoken Words with Applications in Spoken Content Retrieval
- Unveiling the semantic structure of text documents using paragraph-aware Topic Models
- "Is this an example image?" -- Predicting the Relative Abstractness Level of Image and Text
- Dual Attention Network for Product Compatibility and Function Satisfiability Analysis
- Phoneme Level Language Models for Sequence Based Low Resource ASR
- Adaptive additive classification-based loss for deep metric learning
- VideoMCC: a New Benchmark for Video Comprehension
- Using stochastic computation graphs formalism for optimization of sequence-to-sequence model
- Hyperbolic Manifold Regression
- Utilizing FastText for Venue Recommendation
- The Sensitivity of Word Embeddings-based Author Detection Models to Semantic-preserving Adversarial Perturbations
- Alleviating Overfitting for Polysemous Words for Word Representation Estimation Using Lexicons
- Inter-Sense: An Investigation of Sensory Blending in Fiction
- From Algebraic Word Problem to Program: A Formalized Approach
- Patterns versus Characters in Subword-aware Neural Language Modeling
- Converting the Point of View of Messages Spoken to Virtual Assistants
- Controlling the Interaction Between Generation and Inference in Semi-Supervised Variational Autoencoders Using Importance Weighting
- Semantic Preserving Embeddings for Generalized Graphs
- Not just about size - A Study on the Role of Distributed Word Representations in the Analysis of Scientific Publications
- Industry Scale Semi-Supervised Learning for Natural Language Understanding
- Differences between preprints and journal articles : Trial using bioRxiv data
- Learning Embeddings that Capture Spatial Semantics for Indoor Navigation
- Fuzzy paraphrases in learning word representations with a lexicon
- NICT's Corpus Filtering Systems for the WMT18 Parallel Corpus Filtering Task
- Incorporating Relevant Knowledge in Context Modeling and Response Generation
- SocialML: machine learning for social media video creators
- Image Analysis Enhanced Event Detection from Geo-tagged Tweet Streams
- Uncover Sexual Harassment Patterns from Personal Stories by Joint Key Element Extraction and Categorization
- Meta-Embedding as Auxiliary Task Regularization
- Learning Structured Representations of Entity Names using Active Learning and Weak Supervision
- Cross-lingual Word Embeddings beyond Zero-shot Machine Translation
- Improving Word Representations: A Sub-sampled Unigram Distribution for Negative Sampling
- Beyond Next Item Recommendation: Recommending and Evaluating List of Sequences
- Searching for Replacement Classes