# awesome-synthetic-data A curated list of awesome synthetic data tools (open source and commercial). Inspired by [Awesome Synthetic Data](https://github.com/gretelai/awesome-synthetic-data) # Table of content + [Open source tools](#open-source-tools) + [Commercial solutions](#commercial-solutions) + [Online communities](#online-communities) # Open source tools + [Copulas](https://github.com/sdv-dev/Copulas): a Python library for modeling multivariate distributions and sampling from them using copula functions. + [CTGAN](https://github.com/sdv-dev/CTGAN): SDV’s collection of deep learning-based synthetic data generators for single table data. + [DataGene](https://github.com/firmai/datagene): a tool to train, test, and validate datasets, detect and compare dataset similarity between real and synthetic datasets. + [DoppelGANger](https://github.com/fjxmlzn/DoppelGANger): a synthetic data generation framework based on generative adversarial networks (GANs). + [DP_WGAN-UCLANESL](https://github.com/nesl/nist_differential_privacy_synthetic_data_challenge): this solution trains a Wasserstein generative adversarial network (w-GAN) that is trained on the real private dataset. + [DPSyn](https://github.com/usnistgov/PrivacyEngCollabSpace/tree/master/tools/de-identification/Differential-Privacy-Synthetic-Data-Challenge-Algorithms/DPSyn): an algorithm for synthesizing microdata while satisfying differential privacy. + [Faker](https://github.com/joke2k/faker): a Python package that generates fake data (Note: this tool does not generate synthetic data but offers dummy data). + [Generative adversarial nets for synthetic time series data](https://github.com/stefan-jansen/synthetic-data-for-finance): a repository that shows how to create synthetic time-series data using generative adversarial networks (GANs). + [Gretel.ai](https://gretel.ai/): commercial synthetic data vendor that offers open source functionality. + [mirrorGen](https://github.com/DataResponsibly/MirrorDataGenerator): a python tool that generates synthetic data based on user-specified causal relations among features in the data. + [Plait.py](https://github.com/plaitpy/plaitpy): a program for generating fake data from composable yaml templates. + [Pydbgen](https://github.com/tirthajyoti/pydbgen): a Python package that generates a random database table based on the user's choice of data types. + [Smart noise synthesizer](https://smartnoise.org/): a differentially private open source synthesizer for tabular data. + [Synner](https://github.com/huda-lab/synner): an open source tool to generate real-looking synthetic data by visually specifying the properties of the dataset. + [Synth](https://www.getsynth.com/): an open source data-as-code tool that provides a simple CLI workflow for generating consistent data in a scalable way. + [Synthea](https://synthetichealth.github.io/synthea/): an open source synthetic patient generator that models the medical history of synthetic patients. + [Synthetic data vault (SDV)](https://sdv.dev/): one of the first open source synthetic data solutions, SDV provides tools for generating synthetic data for tabular, relational, and time series data. + [TGAN](https://github.com/sdv-dev/TGAN): generative adversarial training for generating synthetic tabular data. + [Tofu](https://github.com/spiros/tofu): a Python library for generating synthetic UK Biobank data. + [Twinify](https://github.com/DPBayes/twinify): a software package for privacy-preserving generation of a synthetic twin to a given sensitive data set. + [YData](https://github.com/ydataai/ydata-synthetic): synthetic structured data generator by YData, a commercial vendor. # Commercial solutions + [Betterdata](https://www.betterdata.ai/): vendor of a privacy-preserving synthetic data solution for AI, data sharing, or product development. + [Datomize](https://www.datomize.com/): vendor of a synthetic data solution for the development, training and testing of AI/ML models, and applications. + [Diveplane](https://diveplane.com/geminai/): vendor of Geminai, a solution to generate synthetic ‘twin’ datasets with the same statistical properties as the original data. + [Facteus](https://www.facteus.com/mimic): vendor of Mimic™ a synthetic data engine to synthesize data assets that protect consumer privacy. + [Gretel](https://gretel.ai/): vendor of a synthetic data generation library and APIs for developers and data practitioners. + [Hazy](https://hazy.com/): vendor of a synthetic data platform for financial institutions that want to conduct data analysis. + [Instill AI](https://instillai.com/): vendor of a solution for synthetic data generation leveraging Generative Adversarial Networks and differential privacy. + [Kymera Labs](https://www.kymera-labs.com/): vendor Synthetic Data Fabrication Software, a solution that generates new data without relying on the ML/GAN approach. + [Mirry.ai](https://www.mirry.ai/main): vendor of a synthetic data platform for generating synthetic data using GANs, available in Community, Cloud or Enterprise editions. + [Mostly AI](https://mostly.ai/): vendor of Mostly Generate, a synthetic data generator that provides as-good-as-real, yet fully anonymous data. + [Replica Analytics](https://replica-analytics.com/): vendor of Replica Synthesis, a software solution that ingests data and builds synthesis models to generate synthetic datasets. + [Sarus technologies](https://www.sarus.tech/): vendor of ML software to help data practitioners leverage sensitive data assets for innovation with privacy guarantees. + [Sogeti](https://www.sogeti.com/services/artificial-intelligence/artificial-data-amplifier/): vendor of Artificial Data Amplifier (ADA), a solution by the Sogeti Testing AI team that generates realistic data based on real data sets. + [Statice](https://www.statice.ai/): vendor of a software solution that generates privacy-preserving synthetic data that can be used as a drop-in replacement for an original dataset. + [Syndata AB](https://syndata.co/): vendor of a synthetic data generator to generate data sets that match the statistical attributes of real data but are entirely synthetic. + [Synthesized](https://www.synthesized.io/): vendor of a DataOps platform enabling data sharing and collaboration across internal groups, remote teams, and external partners. + [Syntheticus](https://syntheticus.ai/): Swiss vendor of a Swiss platform dedicated to generating synthetic data. + [Syntho]([https://www.tonic.ai/](https://www.syntho.ai/)): vendor of AI software for generating synthetic data. + [Tonic](https://www.tonic.ai/): vendor of a synthetic data generator to mimic production data. + [Ydata](https://ydata.ai/): vendor of a synthesizer that mimics statistical information from real data and on new datasets without transforming the original data. # Online communities + [Open SDP](https://opensdp.github.io/data/): an open online community for sharing educational analytic tools and resources. + [OpenSynthetic](https://opensynthetics.com/): an open community for creating and using synthetic data in computer vision and machine learning (ML). + [GenRocket Community](https://community.genrocket.com/) community from GenRocket to ask questions and exchange ideas around test data and synthetic data. + [Synthetic Data Vault Slack channel](https://sdv-space.slack.com/ssb/redirect): the Slack channel from the SDV team.