Modeling Documents with Deep Boltzmann Machines
arXiv:1309.6865
Abstract
We introduce a Deep Boltzmann Machine model suitable for modeling and extracting latent semantic representations from a large unstructured collection of documents. We overcome the apparent difficulty of training a DBM with judicious parameter tying. This parameter tying enables an efficient pretraining algorithm and a state initialization scheme that aids inference. The model can be trained just as efficiently as a standard Restricted Boltzmann Machine. Our experiments show that the model assigns better log probability to unseen data than the Replicated Softmax model. Features extracted from our model outperform LDA, Replicated Softmax, and DocNADE models on document retrieval and document classification tasks.
Appears in Proceedings of the Twenty-Ninth Conference on Uncertainty in Artificial Intelligence (UAI2013)
References in corpus (1)
Cited by in corpus (18)
- Distributed Representations of Sentences and Documents
- Deep Latent Dirichlet Allocation with Topic-Layer-Adaptive Stochastic Gradient Riemannian MCMC
- Combining LSTM and Latent Topic Modeling for Mortality Prediction
- Neural Simpletrons - Minimalistic Directed Generative Networks for Learning with Few Labels
- Learning Document Embeddings by Predicting N-grams for Sentiment Classification of Long Movie Reviews
- Document Neural Autoregressive Distribution Estimation
- Semantic Regularities in Document Representations
- Conditional Restricted Boltzmann Machines for Cold Start Recommendations
- Ordering-sensitive and Semantic-aware Topic Modeling
- Zero-Shot Clinical Acronym Expansion via Latent Meaning Cells
- Restricted Boltzmann Machine and Deep Belief Network: Tutorial and Survey
- Restricted Boltzmann Machines with Gaussian Visible Units Guided by Pairwise Constraints
- Modeling correlations in spontaneous activity of visual cortex with centered Gaussian-binary deep Boltzmann machines
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- Understanding the Behaviour of the Empirical Cross-Entropy Beyond the Training Distribution
- Adaptive Bayesian Sampling with Monte Carlo EM
- Efficient Learning for Undirected Topic Models
- Learning Topics using Semantic Locality