Probabilistic Latent Semantic Analysis
arXiv:1301.6705
Abstract
Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which stems from linear algebra and performs a Singular Value Decomposition of co-occurrence tables, the proposed method is based on a mixture decomposition derived from a latent class model. This results in a more principled approach which has a solid foundation in statistics. In order to avoid overfitting, we propose a widely applicable generalization of maximum likelihood model fitting by tempered EM. Our approach yields substantial and consistent improvements over Latent Semantic Analysis in a number of experiments.
Appears in Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence (UAI1999)
Cited by in corpus (12)
- Probabilistic Models for Unified Collaborative and Content-Based Recommendation in Sparse-Data Environments
- Visualizing Topics with Multi-Word Expressions
- Expectation-Propogation for the Generative Aspect Model
- Incorporating Side Information in Probabilistic Matrix Factorization with Gaussian Processes
- Mixed membership analysis of high-throughput interaction studies: Relational data
- A Tutorial on Probabilistic Latent Semantic Analysis
- Using Access Data for Paper Recommendations on ArXiv.org
- Factorized Topic Models
- Message-Passing Inference on a Factor Graph for Collaborative Filtering
- Approximate Maximum A Posteriori Inference with Entropic Priors
- Comparison Latent Semantic and WordNet Approach for Semantic Similarity Calculation
- The Latent Bernoulli-Gauss Model for Data Analysis