2 papers
cs.CL2026
Determinants of Training Corpus Size for Clinical Text Classification
Jaya Chaturvedi, Saniya Deshpande, Chenkai Ma +5
Introduction: Clinical text classification using natural language processing (NLP) models requires adequate training data to achieve optimal performance. For that, 200-500 document…
cs.LG2023
Sample Size in Natural Language Processing within Healthcare Research
Jaya Chaturvedi, Diana Shamsutdinova, Felix Zimmer +4
Sample size calculation is an essential step in most data-based disciplines. Large enough samples ensure representativeness of the population and determine the precision of estimat…