Experimental Standards for Deep Learning in Natural Language Processing Research
arXiv:2204.06251
Abstract
The field of Deep Learning (DL) has undergone explosive growth during the last decade, with a substantial impact on Natural Language Processing (NLP) as well. Yet, compared to more established disciplines, a lack of common experimental standards remains an open challenge to the field at large. Starting from fundamental scientific principles, we distill ongoing discussions on experimental standards in NLP into a single, widely-applicable methodology. Following these best practices is crucial to strengthen experimental evidence, improve reproducibility and support scientific progress. These standards are further collected in a public repository to help them transparently adapt to future needs.
References in corpus (11)
- Practical Bayesian Optimization of Machine Learning Algorithms
- Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
- Decision Transformer: Reinforcement Learning via Sequence Modeling
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
- Compute Trends Across Three Eras of Machine Learning
- A Systematic Review of Reproducibility Research in Natural Language Processing
- Nonparametric Estimation of Heterogeneous Treatment Effects: From Theory to Learning Algorithms
- deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks
- Best Practices for Managing Data Annotation Projects
- Deep Learning Reproducibility and Explainable AI (XAI)