Synthcity: facilitating innovative use cases of synthetic data in different data modalities
arXiv:2301.07573
Abstract
Synthcity is an open-source software package for innovative use cases of synthetic data in ML fairness, privacy and augmentation across diverse tabular data modalities, including static data, regular and irregular time series, data with censoring, multi-source data, composite data, and more. Synthcity provides the practitioners with a single access point to cutting edge research and tools in synthetic data. It also offers the community a playground for rapid experimentation and prototyping, a one-stop-shop for SOTA benchmarks, and an opportunity for extending research impact. The library can be accessed on GitHub (https://github.com/vanderschaarlab/synthcity) and pip (https://pypi.org/project/synthcity/). We warmly invite the community to join the development effort by providing feedback, reporting bugs, and contributing code.
Cited by in corpus (6)
- SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
- Synthetic Data: Revisiting the Privacy-Utility Trade-off
- DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
- To democratize research with sensitive data, we should make synthetic data more accessible
- Methods for generating and evaluating synthetic longitudinal patient data: a systematic review
- Synthetic Tabular Data: Methods, Attacks and Defenses