Semi-supervised Sequence Learning
arXiv:1511.01432
Abstract
We present two approaches that use unlabeled data to improve sequence learning with recurrent networks. The first approach is to predict what comes next in a sequence, which is a conventional language model in natural language processing. The second approach is to use a sequence autoencoder, which reads the input sequence into a vector and predicts the input sequence again. These two algorithms can be used as a "pretraining" step for a later supervised sequence learning algorithm. In other words, the parameters obtained from the unsupervised step can be used as a starting point for other supervised training models. In our experiments, we find that long short term memory recurrent networks after being pretrained with the two approaches are more stable and generalize better. With pretraining, we are able to train long short term memory recurrent networks up to a few hundred timesteps, thereby achieving strong performance in many text classification tasks, such as IMDB, DBpedia and 20 Newsgroups.
References in corpus (9)
- Sequence to Sequence Learning with Neural Networks
- LSTM: A Search Space Odyssey
- Distributed Representations of Sentences and Documents
- Recurrent Neural Network Regularization
- Unsupervised Learning of Video Representations using LSTMs
- A Neural Conversational Model
- Convolutional Neural Networks for Sentence Classification
- Listen, Attend and Spell
- On Using Very Large Target Vocabulary for Neural Machine Translation
Cited by in corpus (167)
- Learning Transferable Visual Models From Natural Language Supervision
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Language Models are Few-Shot Learners
- On the Opportunities and Risks of Foundation Models
- Longformer: The Long-Document Transformer
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- Unsupervised Data Augmentation for Consistency Training
- Pre-trained Models for Natural Language Processing: A Survey
- Evaluating Large Language Models Trained on Code
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- TAPAS: Weakly Supervised Table Parsing via Pre-training
- REALM: Retrieval-Augmented Language Model Pre-Training
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Fine-Tuning Language Models from Human Preferences
- Rethinking Pre-training and Self-training
- Deep Learning Based Text Classification: A Comprehensive Review
- BARTScore: Evaluating Generated Text as Text Generation
- A Simple Method for Commonsense Reasoning
- A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music
- A Transformer-based approach to Irony and Sarcasm detection
- Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language
- Realistic Evaluation of Deep Semi-Supervised Learning Algorithms
- Incorporating BERT into Neural Machine Translation
- Supervised Multimodal Bitransformers for Classifying Images and Text
- Unsupervised Learning via Meta-Learning
- Learning and Evaluating General Linguistic Intelligence
- Parameter-Efficient Transfer Learning for NLP
- ERNIE: Enhanced Language Representation with Informative Entities
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Analysing Lexical Semantic Change with Contextualised Word Representations
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach
- Exploring the Limits of Out-of-Distribution Detection
- How to Fine-Tune BERT for Text Classification?
- Vector-quantized Image Modeling with Improved VQGAN
- A Mutual Information Maximization Perspective of Language Representation Learning
- Selfie: Self-supervised Pretraining for Image Embedding
- Finetuned Language Models Are Zero-Shot Learners
- Sample Efficient Text Summarization Using a Single Pre-Trained Transformer
- FlauBERT: Unsupervised Language Model Pre-training for French
- Deep Learning-based Sentiment Classification: A Comparative Survey
- Weight Poisoning Attacks on Pre-trained Models
- Gmail Smart Compose: Real-Time Assisted Writing
- Reweighted Proximal Pruning for Large-Scale Language Representation
- Context2Name: A Deep Learning-Based Approach to Infer Natural Variable Names from Usage Contexts
- Grad-SAM: Explaining Transformers via Gradient Self-Attention Maps
- Uncertainty-aware Self-training for Text Classification with Few Labels
- Language Modeling Teaches You More Syntax than Translation Does: Lessons Learned Through Auxiliary Task Analysis
- Neural ranking models for document retrieval
- Efficient Adaptation of Pretrained Transformers for Abstractive Summarization
- Attention-based Clinical Note Summarization
- Similarity Learning for Authorship Verification in Social Media
- Learning to summarize from human feedback
- Deep-Sentiment: Sentiment Analysis Using Ensemble of CNN and Bi-LSTM Models
- AP-10K: A Benchmark for Animal Pose Estimation in the Wild
- Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
- A Multimodal Memes Classification: A Survey and Open Research Issues
- Program Synthesis with Large Language Models
- Pre-trained Language Model Representations for Language Generation
- Parameter-Efficient Transfer Learning with Diff Pruning
- Learning Cross-Context Entity Representations from Text
- Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
- Improving Readability for Automatic Speech Recognition Transcription
- MKD: a Multi-Task Knowledge Distillation Approach for Pretrained Language Models
- An Embarrassingly Simple Approach for Transfer Learning from Pretrained Language Models
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- Using Similarity Measures to Select Pretraining Data for NER
- A Scalable Framework for Multilevel Streaming Data Analytics using Deep Learning
- Improving Robustness of Task Oriented Dialog Systems
- On Linear Identifiability of Learned Representations
- Can Unconditional Language Models Recover Arbitrary Sentences?
- Financial Aspect-Based Sentiment Analysis using Deep Representations
- Temporal-Clustering Invariance in Irregular Healthcare Time Series
- ERICA: Improving Entity and Relation Understanding for Pre-trained Language Models via Contrastive Learning
- Clinical Relation Extraction Using Transformer-based Models
- Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes
- Primer: Searching for Efficient Transformers for Language Modeling
- Knowledge Enhanced Attention for Robust Natural Language Inference
- Exploiting Deep Learning for Persian Sentiment Analysis
- Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space
- Interpretable Adversarial Training for Text
- State-Regularized Recurrent Neural Networks
- Neural Semi-supervised Learning for Text Classification Under Large-Scale Pretraining
- Learning Compressed Sentence Representations for On-Device Text Processing
- indicnlp@kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages
- Towards Domain-Agnostic Contrastive Learning
- A Financial Service Chatbot based on Deep Bidirectional Transformers
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
- Multi-task Learning based Pre-trained Language Model for Code Completion
- Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs
- The Rich Get Richer: Disparate Impact of Semi-Supervised Learning
- Unsupervised Domain Adaptation of Contextualized Embeddings for Sequence Labeling
- Independent language modeling architecture for end-to-end ASR
- Fast CRDNN: Towards on Site Training of Mobile Construction Machines
- Trimming and Improving Skip-thought Vectors
- Surprisal-Triggered Conditional Computation with Neural Networks
- A Survey On Neural Word Embeddings
- Unsupervised Extractive Summarization by Pre-training Hierarchical Transformers
- UER: An Open-Source Toolkit for Pre-training Models
- mu-Forcing: Training Variational Recurrent Autoencoders for Text Generation
- Contextual Lensing of Universal Sentence Representations
- A Survey on Understanding, Visualizations, and Explanation of Deep Neural Networks
- Denoising based Sequence-to-Sequence Pre-training for Text Generation
- Target-Embedding Autoencoders for Supervised Representation Learning
- Leveraging Just a Few Keywords for Fine-Grained Aspect Detection Through Weakly Supervised Co-Training
- Enhancing Relation Extraction Using Syntactic Indicators and Sentential Contexts
- Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation
- Improving BERT Model Using Contrastive Learning for Biomedical Relation Extraction
- Table2answer: Read the database and answer without SQL
- mmFall: Fall Detection using 4D MmWave Radar and a Hybrid Variational RNN AutoEncoder
- K-XLNet: A General Method for Combining Explicit Knowledge with Language Model Pretraining
- Repurposing Decoder-Transformer Language Models for Abstractive Summarization
- Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks
- Neural Attentive Bag-of-Entities Model for Text Classification
- Bridging the Knowledge Gap: Enhancing Question Answering with World and Domain Knowledge
- Inferring Offensiveness In Images From Natural Language Supervision
- Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages
- A Practical Framework for Relation Extraction with Noisy Labels Based on Doubly Transitional Loss
- Weakly-Supervised Hierarchical Models for Predicting Persuasive Strategies in Good-faith Textual Requests
- Investigation of BERT Model on Biomedical Relation Extraction Based on Revised Fine-tuning Mechanism
- Instrument Classification of Solo Sheet Music Images
- Masked ELMo: An evolution of ELMo towards fully contextual RNN language models
- DETECT: Deep Trajectory Clustering for Mobility-Behavior Analysis
- Information Extraction from Scientific Literature for Method Recommendation
- Non-local Recurrent Neural Memory for Supervised Sequence Modeling
- Performance Analysis of Deep Learning Workloads on Leading-edge Systems
- Learning to Explicitate Connectives with Seq2Seq Network for Implicit Discourse Relation Classification
- An approach based on Combination of Features for automatic news retrieval
- Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project
- Fast and Accurate Deep Bidirectional Language Representations for Unsupervised Learning
- Typing Errors in Factual Knowledge Graphs: Severity and Possible Ways Out
- Personalizing Pre-trained Models
- Multi-Scale One-Class Recurrent Neural Networks for Discrete Event Sequence Anomaly Detection
- A Neural Attention Model for Categorizing Patient Safety Events
- Learning Temporal Action Proposals With Fewer Labels
- BERTweetFR : Domain Adaptation of Pre-Trained Language Models for French Tweets
- Multirate Training of Neural Networks
- BERT_SE: A Pre-trained Language Representation Model for Software Engineering
- Bounding the expected run-time of nonconvex optimization with early stopping
- Towards Controlled Transformation of Sentiment in Sentences
- Learning Dynamic BERT via Trainable Gate Variables and a Bi-modal Regularizer
- Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders
- Interpretable Structure-aware Document Encoders with Hierarchical Attention
- Indirect Identification of Psychosocial Risks from Natural Language
- Metric Learning for Dynamic Text Classification
- Pre-train, Interact, Fine-tune: A Novel Interaction Representation for Text Classification
- SNU_IDS at SemEval-2018 Task 12: Sentence Encoder with Contextualized Vectors for Argument Reasoning Comprehension
- Distributionally Robust Language Modeling
- Multimodal Story Generation on Plural Images
- Eliciting Knowledge from Experts:Automatic Transcript Parsing for Cognitive Task Analysis
- Pre-Trained Models for Heterogeneous Information Networks
- Public Health Informatics: Proposing Causal Sequence of Death Using Neural Machine Translation
- Deep Learning Meets Projective Clustering
- It's FLAN time! Summing feature-wise latent representations for interpretability
- Toward Human-Level Artificial Intelligence
- Deep Symbolic Representation Learning for Heterogeneous Time-series Classification
- Semi-Supervised Self-Growing Generative Adversarial Networks for Image Recognition
- Recognising Biomedical Names: Challenges and Solutions
- Knowledge Distillation from BERT Transformer to Speech Transformer for Intent Classification
- Pre-Training Transformers as Energy-Based Cloze Models
- A More Efficient Chinese Named Entity Recognition base on BERT and Syntactic Analysis