Learned in Translation: Contextualized Word Vectors
arXiv:1708.00107
Abstract
Computer vision has benefited from initializing multiple deep layers with weights pretrained on large supervised training sets like ImageNet. Natural language processing (NLP) typically sees initialization of only the lowest layer of deep models with pretrained word vectors. In this paper, we use a deep LSTM encoder from an attentional sequence-to-sequence model trained for machine translation (MT) to contextualize word vectors. We show that adding these context vectors (CoVe) improves performance over using only unsupervised word and character vectors on a wide variety of common NLP tasks: sentiment analysis (SST, IMDb), question classification (TREC), entailment (SNLI), and question answering (SQuAD). For fine-grained sentiment analysis and entailment, CoVe improves performance of our baseline models to the state of the art.
Cited by in corpus (55)
- Pre-trained Models for Natural Language Processing: A Survey
- The Natural Language Decathlon: Multitask Learning as Question Answering
- A Transformer-based approach to Irony and Sarcasm detection
- Parameter-Efficient Transfer Learning for NLP
- Backdoor Embedding in Convolutional Neural Network Models via Invisible Perturbation
- Zero-Shot Cross-lingual Classification Using Multilingual Neural Machine Translation
- Dynamic Integration of Background Knowledge in Neural NLU Systems
- Language Modeling Teaches You More Syntax than Translation Does: Lessons Learned Through Auxiliary Task Analysis
- Semantic Sentence Matching with Densely-connected Recurrent and Co-attentive Information
- Efficient Adaptation of Pretrained Transformers for Abstractive Summarization
- Fully Transformer Networks for Semantic Image Segmentation
- Efficient and Robust Question Answering from Minimal Context over Documents
- Pseudo-task Augmentation: From Deep Multitask Learning to Intratask Sharing---and Back
- Socially Enhanced Situation Awareness from Microblogs using Artificial Intelligence: A Survey
- Semi-Supervised Sequence Modeling with Cross-View Training
- Extracting Multiple-Relations in One-Pass with Pre-Trained Transformers
- Lessons from Natural Language Inference in the Clinical Domain
- Pre-Trained Models: Past, Present and Future
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Interpreting Recurrent and Attention-Based Neural Models: a Case Study on Natural Language Inference
- Transfer Learning in Multilingual Neural Machine Translation with Dynamic Vocabulary
- Few-Shot Learning as Domain Adaptation: Algorithm and Analysis
- Investigating the Effectiveness of Representations Based on Word-Embeddings in Active Learning for Labelling Text Datasets
- Evaluation of taxonomic and neural embedding methods for calculating semantic similarity
- Diverse Pretrained Context Encodings Improve Document Translation
- Multi-task Learning for Universal Sentence Embeddings: A Thorough Evaluation using Transfer and Auxiliary Tasks
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Sogou Machine Reading Comprehension Toolkit
- Improving AMR Parsing with Sequence-to-Sequence Pre-training
- Mark my Word: A Sequence-to-Sequence Approach to Definition Modeling
- Multi-view Sentence Representation Learning
- There is no Artificial General Intelligence
- A Survey On Neural Word Embeddings
- Sense representations for Portuguese: experiments with sense embeddings and deep neural language models
- Learning Robust, Transferable Sentence Representations for Text Classification
- Leveraging Advantages of Interactive and Non-Interactive Models for Vector-Based Cross-Lingual Information Retrieval
- Pedestrian Detection with Autoregressive Network Phases
- Improving Sentence Representations with Consensus Maximisation
- Still a Pain in the Neck: Evaluating Text Representations on Lexical Composition
- Learning to Embed Sentences Using Attentive Recursive Trees
- WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse
- Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework
- BERT_SE: A Pre-trained Language Representation Model for Software Engineering
- Dynamic Compositionality in Recursive Neural Networks with Structure-aware Tag Representations
- AWE: Asymmetric Word Embedding for Textual Entailment
- Enhance Long Text Understanding via Distilled Gist Detector from Abstractive Summarization
- Contextual Salience for Fast and Accurate Sentence Vectors
- LV-BERT: Exploiting Layer Variety for BERT
- SNU_IDS at SemEval-2018 Task 12: Sentence Encoder with Contextualized Vectors for Argument Reasoning Comprehension
- Towards Controlled Transformation of Sentiment in Sentences
- Pre-train, Interact, Fine-tune: A Novel Interaction Representation for Text Classification
- Learning to Compose over Tree Structures via POS Tags
- Re-Evaluating GermEval17 Using German Pre-Trained Language Models
- Multi-level Head-wise Match and Aggregation in Transformer for Textual Sequence Matching
- Towards Language Agnostic Universal Representations