3 papers
math.ST2026
Data Reuse and the Long Shadow of Error: Splitting, Subsampling, and Prospectively Managing Inferential Errors
Reid Dale, Jordan Rodu, Maria E. Currie +1
When multiple investigators analyze a common dataset, the data reuse induces dependence across testing procedures, affecting the distribution of errors. Existing techniques of mana…
stat.ME2026
Synthetic Data, Information, and Prior Knowledge: Why Synthetic Data Augmentation to Boost Sample Doesn't Work for Statistical Inference
Reid Dale, Jordan Rodu, Mike Baiocchi
The use of synthetic data to deidentify data and to improve predictive models is well-attested to. The augmentation of datasets using synthetically generated data is an alluring pr…
math.ST2025
Data Gluttony: Epistemic Risks, Dependent Testing and Data Reuse in Large Datasets
Reid Dale, Jordan Rodu, Maria E. Currie +1
Large-scale registries have collected vast amounts of data which has enabled investigators to efficiently conduct studies of observational data. Common practice is for investigator…