Generative and Discriminative Text Classification with Recurrent Neural Networks
arXiv:1703.01898
Abstract
We empirically characterize the performance of discriminative and generative LSTM models for text classification. We find that although RNN-based generative models are more powerful than their bag-of-words ancestors (e.g., they account for conditional dependencies across words in a document), they have higher asymptotic error rates than discriminatively trained RNN models. However we also find that generative models approach their asymptotic error rate more rapidly than their discriminative counterparts---the same pattern that Ng & Jordan (2001) proved holds for linear classification models that make more naive conditional independence assumptions. Building on this finding, we hypothesize that RNN-based generative classification models will be more robust to shifts in the data distribution. This hypothesis is confirmed in a series of experiments in zero-shot and continual learning settings that show that generative models substantially outperform discriminative models.
References in corpus (1)
Cited by in corpus (26)
- Patient2Vec: A Personalized Interpretable Deep Representation of the Longitudinal Electronic Health Record
- How to Fine-Tune BERT for Text Classification?
- Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach
- Hash Embeddings for Efficient Word Representations
- Description Based Text Classification with Reinforcement Learning
- Learning to Remember More with Less Memorization
- Identifying emergency stages in Facebook posts of police departments with convolutional and recurrent neural networks and support vector machines
- Medical Diagnosis From Laboratory Tests by Combining Generative and Discriminative Learning
- Hierarchical CVAE for Fine-Grained Hate Speech Classification
- GeDi: Generative Discriminator Guided Sequence Generation
- Low-Rank RNN Adaptation for Context-Aware Language Modeling
- Benchmark Performance of Machine And Deep Learning Based Methodologies for Urdu Text Document Classification
- Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection In Task Oriented Dialog
- Unsupervised Label Refinement Improves Dataless Text Classification
- Memory and attention in deep learning
- Latent-Variable Generative Models for Data-Efficient Text Classification
- Large-Scale Spectrum Occupancy Learning via Tensor Decomposition and LSTM Networks
- Adaptive Region Embedding for Text Classification
- NatCat: Weakly Supervised Text Classification with Naturally Annotated Resources
- Know thy corpus! Robust methods for digital curation of Web corpora
- CrowdTSC: Crowd-based Neural Networks for Text Sentiment Classification
- Cascaded Semantic and Positional Self-Attention Network for Document Classification
- Energy-based Unknown Intent Detection with Data Manipulation
- Maps Search Misspelling Detection Leveraging Domain-Augmented Contextual Representations
- Generative and Discriminative Deep Belief Network Classifiers: Comparisons Under an Approximate Computing Framework
- Discriminatively-Tuned Generative Classifiers for Robust Natural Language Inference