Experimental Standards for Deep Learning in Natural Language Processing Research
arXiv:2204.06251
Abstract
The field of Deep Learning (DL) has undergone explosive growth during the last decade, with a substantial impact on Natural Language Processing (NLP) as well. Yet, compared to more established disciplines, a lack of common experimental standards remains an open challenge to the field at large. Starting from fundamental scientific principles, we distill ongoing discussions on experimental standards in NLP into a single, widely-applicable methodology. Following these best practices is crucial to strengthen experimental evidence, improve reproducibility and support scientific progress. These standards are further collected in a public repository to help them transparently adapt to future needs.
References in corpus (16)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Practical Bayesian Optimization of Machine Learning Algorithms
- Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
- Decision Transformer: Reinforcement Learning via Sequence Modeling
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
- Compute Trends Across Three Eras of Machine Learning
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- A Systematic Review of Reproducibility Research in Natural Language Processing
- Understanding the Failure Modes of Out-of-Distribution Generalization
- Hyperparameter Optimization: Foundations, Algorithms, Best Practices and Open Challenges
- Nonparametric Estimation of Heterogeneous Treatment Effects: From Theory to Learning Algorithms
- deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks
- Best Practices for Managing Data Annotation Projects
- Deep Learning Reproducibility and Explainable AI (XAI)