Optimal Ratio for Data Splitting
arXiv:2202.03326 · doi:10.1002/sam.11583
Abstract
It is common to split a dataset into training and testing sets before fitting a statistical or machine learning model. However, there is no clear guidance on how much data should be used for training and testing. In this article we show that the optimal splitting ratio is , where is the number of parameters in a linear regression model that explains the data well.
References in corpus (2)
Cited by in corpus (8)
- Data Readiness for AI: A 360-Degree Survey
- Reconstructing complex states of a 20-qubit quantum simulator
- On Real-time Image Reconstruction with Neural Networks for MRI-guided Radiotherapy
- Neural Network Emulation of Flow in Heavy-Ion Collisions at Intermediate Energies
- On the Cost of Model-Serving Frameworks: An Experimental Evaluation
- Discovery of Fatigue Strength Models via Feature Engineering and automated eXplainable Machine Learning applied to the welded Transverse Stiffener
- Evolutions of in-medium baryon-baryon scattering cross sections and stiffness of dense nuclear matter from Bayesian analyses of FOPI proton flow excitation functions
- Machine Learning-Driven Insights into Excitonic Effects in 2D Materials