Easy over Hard: A Case Study on Deep Learning
arXiv:1703.00133 · doi:10.1145/3106237.3106256
Abstract
While deep learning is an exciting new technique, the benefits of this method need to be assessed with respect to its computational cost. This is particularly important for deep learning since these learners need hours (to weeks) to train the model. Such long training time limits the ability of (a)~a researcher to test the stability of their conclusion via repeated runs with different random seeds; and (b)~other researchers to repeat, improve, or even refute that original work. For example, recently, deep learning was used to find which questions in the Stack Overflow programmer discussion forum can be linked together. That deep learning system took 14 hours to execute. We show here that applying a very simple optimizer called DE to fine tune SVM, it can achieve similar (and sometimes better) results. The DE approach terminated in 10 minutes; i.e. 84 times faster hours than deep learning method. We offer these results as a cautionary tale to the software analytics community and suggest that not every new innovation should be applied without critical analysis. If researchers deploy some new and expensive process, that work should be baselined against some simpler and faster alternatives.
12 pages, 6 figures, accepted at FSE2017
References in corpus (3)
Cited by in corpus (21)
- What is Wrong with Topic Modeling? (and How to Fix it Using Search-based Software Engineering)
- Is "Better Data" Better than "Better Data Miners"? (On the Benefits of Tuning SMOTE for Defect Prediction)
- No More Fine-Tuning? An Experimental Evaluation of Prompt Tuning in Code Intelligence
- Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
- Opportunities and Challenges in Code Search Tools
- 500+ Times Faster Than Deep Learning (A Case Study Exploring Faster Methods for Text Mining StackOverflow)
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?
- Cross-Domain Deep Code Search with Meta Learning
- Duplicate Bug Report Detection: How Far Are We?
- Better Software Analytics via "DUO": Data Mining Algorithms Using/Used-by Optimizers
- ArguLens: Anatomy of Community Opinions On Usability Issues Using Argumentation Models
- Semantic Source Code Models Using Identifier Embeddings
- On Using Machine Learning to Identify Knowledge in API Reference Documentation
- Data-Driven Search-based Software Engineering
- Is One Hyperparameter Optimizer Enough?
- Feature Sets in Just-in-Time Defect Prediction: An Empirical Evaluation
- A Machine Learning Based Ensemble Method for Automatic Multiclass Classification of Decisions
- Predicting Breakdowns in Cloud Services (with SPIKE)
- ARCLIN: Automated API Mention Resolution for Unformatted Texts
- An Empirical Study on Code Review Activity Prediction and Its Impact in Practice
- AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding