How to avoid machine learning pitfalls: a guide for academic researchers
arXiv:2108.02497 · doi:10.1016/j.patter.2024.101046
Abstract
Mistakes in machine learning practice are commonplace, and can result in a loss of confidence in the findings and products of machine learning. This guide outlines common mistakes that occur when using machine learning, and what can be done to avoid them. Whilst it should be accessible to anyone with a basic understanding of machine learning techniques, it focuses on issues that are of particular concern within academic research, such as the need to do rigorous comparisons and reach valid conclusions. It covers five stages of the machine learning process: what to do before model building, how to reliably build models, how to robustly evaluate models, how to compare models fairly, and how to report results.
References in corpus (9)
- Deep Learning in Neural Networks: An Overview
- Recent Trends in the Use of Statistical Tests for Comparing Swarm and Evolutionary Computing Algorithms: Practical Guidelines and a Critical Review
- Data Augmentation techniques in time series domain: A survey and taxonomy
- Ablation Studies in Artificial Neural Networks
- A Hierarchy of Limitations in Machine Learning
- Operationalizing Machine Learning: An Interview Study
- Privacy in Large Language Models: Attacks, Defenses and Future Directions
- A Comprehensive Survey on Data Augmentation
- Data Cleaning and Machine Learning: A Systematic Literature Review
Cited by in corpus (5)
- Map-Based Path Loss Prediction in Multiple Cities Using Convolutional Neural Networks
- Reassessing feature-based Android malware detection in a contemporary context
- A validity-guided workflow for robust large language model research in psychology
- From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
- Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers