What is Wrong with Topic Modeling? (and How to Fix it Using Search-based Software Engineering)
arXiv:1608.08176 · doi:10.1016/j.infsof.2018.02.005
Abstract
Context: Topic modeling finds human-readable structures in unstructured textual data. A widely used topic modeler is Latent Dirichlet allocation. When run on different datasets, LDA suffers from "order effects" i.e. different topics are generated if the order of training data is shuffled. Such order effects introduce a systematic error for any study. This error can relate to misleading results;specifically, inaccurate topic descriptions and a reduction in the efficacy of text mining classification results. Objective: To provide a method in which distributions generated by LDA are more stable and can be used for further analysis. Method: We use LDADE, a search-based software engineering tool that tunes LDA's parameters using DE (Differential Evolution). LDADE is evaluated on data from a programmer information exchange site (Stackoverflow), title and abstract text of thousands ofSoftware Engineering (SE) papers, and software defect reports from NASA. Results were collected across different implementations of LDA (Python+Scikit-Learn, Scala+Spark); across different platforms (Linux, Macintosh) and for different kinds of LDAs (VEM,or using Gibbs sampling). Results were scored via topic stability and text mining classification accuracy. Results: In all treatments: (i) standard LDA exhibits very large topic instability; (ii) LDADE's tunings dramatically reduce cluster instability; (iii) LDADE also leads to improved performances for supervised as well as unsupervised learning. Conclusion: Due to topic instability, using standard LDA with its "off-the-shelf" settings should now be depreciated. Also, in future, we should require SE papers that use LDA to test and (if needed) mitigate LDA topic instability. Finally, LDADE is a candidate technology for effectively and efficiently reducing that instability.
15 pages + 2 page references. Accepted to IST
References in corpus (6)
- Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation
- NLTK: The Natural Language Toolkit
- MLlib: Machine Learning in Apache Spark
- On Smoothing and Inference for Topic Models
- Tuning for Software Analytics: is it Really Necessary?
- Easy over Hard: A Case Study on Deep Learning
Cited by in corpus (34)
- Comparison of Topic Modelling Approaches in the Banking Context
- Measuring LDA Topic Stability from Clusters of Replicated Runs
- Challenges in Docker Development: A Large-scale Study Using Stack Overflow
- How to "DODGE" Complex Software Analytics?
- Finding Trends in Software Research
- Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors
- Why is Differential Evolution Better than Grid Search for Tuning Defect Predictors?
- Better Software Analytics via "DUO": Data Mining Algorithms Using/Used-by Optimizers
- Hyperparameter Optimization for Effort Estimation
- Better Security Bug Report Classification via Hyperparameter Optimization
- ArguLens: Anatomy of Community Opinions On Usability Issues Using Argumentation Models
- Semantic Source Code Models Using Identifier Embeddings
- Revisiting Process versus Product Metrics: a Large Scale Analysis
- Simpler Hyperparameter Optimization for Software Analytics: Why, How, When?
- Is One Hyperparameter Optimizer Enough?
- Active Learning for Identifying Disaster-Related Tweets: A Comparison with Keyword Filtering and Generic Fine-Tuning
- Improving Reliability of Latent Dirichlet Allocation by Assessing Its Stability Using Clustering Techniques on Replicated Runs
- Topic Modelling of Empirical Text Corpora: Validity, Reliability, and Reproducibility in Comparison to Semantic Maps
- Determination of the Number of Topics Intrinsically: Is It Possible?
- Profiling Software Developers with Process Mining and N-Gram Language Models
- Analysis and Detection of Information Types of Open Source Software Issue Discussions
- Why Software Effort Estimation Needs SBSE
- Preference Discovery in Large Product Lines
- Defect Reduction Planning (using TimeLIME)
- Better Data Labelling with EMBLEM (and how that Impacts Defect Prediction)
- Model Review: A PROMISEing Opportunity
- Applications of Psychological Science for Actionable Analytics
- Can You Explain That, Better? Comprehensible Text Analytics for SE Applications
- Semiparametric Latent Topic Modeling on Consumer-Generated Corpora
- Bootstrapping Cookbooks for APIs from Crowd Knowledge on Stack Overflow
- How to Better Distinguish Security Bug Reports (using Dual Hyperparameter Optimization
- Mining Scientific Workflows for Anomalous Data Transfers
- Modeling Hierarchical Usage Context for Software Exceptions based on Interaction Data
- Using meaning instead of words to track topics