Synthia: multidimensional synthetic data generation in Python
The Journal of Open Source Software, vol. 6, pp. 2863
Abstract
NeedSynthetic data -artificially generated data that mimic the original (observed) data by preserving relationships between variables (Nowok et al., 2016) -may be useful in several areas such as healthcare, finance, data science, and machine learning (Dahmen & Cook, 2019;Kamthe et al., 2021;Nowok et al., 2016;Patki et al., 2016).As such, copula-based data generation models -probabilistic models that allow for the statistical properties of observed data to be modelled in terms of individual behavior and (inter-)dependencies (Joe, 2014) -have shown potential in several applications such as finance, data science, and meteorology (Kamthe et al., 2021;Li et al., 2020;Meyer, Nagler, et al., 2021;Patki et al., 2016).Although copula-based data generation tools have been developed for tabular data -e.g. the Synthetic Data Vault project using Gaussian copulas and generative adversarial networks (Patki et al., 2016;Xu & Veeramachaneni, 2018), or the Synthetic Data Generation via Gaussian Copula (Li et al., 2020) -in computational sciences such as weather and climate, data often consist of large, labelled multidimensional datasets with complex dependencies.
Authors 1
-
Affiliation as printed
Mathematical Institute , Leiden University , Leiden , The Netherlands
Mathematical Institute, Leiden University, Leiden, The Netherlands
Cited by 26 stored of 27
References 17
-
W390686421details pending0citations
-
W2902901670details pending0citations
-
W2921914364details pending0citations
-
W3120369609details pending0citations
-
W3129829742details pending0citations
-
W3204034810details pending0citations
-
W6931338223details pending0citations
-
W6950250550details pending0citations
17 results