Generate synthetic datasets from metadata collected from the public CBS (Statistics Netherlands) microdata catalog.
Warning
Experimental prototype. This project is a work in progress, not an official CBS product, and is not intended for production or policy decisions. Generated data are not validated as statistically representative of real microdata and have no formal privacy or disclosure-control guarantee. Review the code and outputs for your use case.
This experiment explores how far public catalog metadata can drive the generation of synthetic, structurally plausible datasets, including datasets linked by shared keys. The aim is to provide reproducible fixtures for developing and testing data workflows when real records are unavailable, while learning where metadata-driven generation falls short. It does not aim to reproduce real population distributions or certify that a workflow will work correctly on real microdata.
This project reads the local SQLite database built by the separate microdata_catalogus project to:
- discover dataset structures and variable definitions
- generate synthetic values that respect metadata constraints
- produce linked datasets with referential integrity
The current prototype includes experiments for:
- all 7 key types (person, business, job, household, object, address, education)
- composite key-aware linking (for example
RINPERSOON + RINPERSOONSsampled as a pair) - metadata JSON integration (
code_list,usage_notes,length) - case-insensitive dataset and key matching
- period-aware date generation and start/end ordering constraints
The generated catalog database is not included in the microdata_catalogus repository. Install uv and Python 3.12 or later,
then clone that repository next to microdata_synth and build the database before running the generator:
# From the microdata_synth repository root (skip cloning if it is already there)
git clone https://github.com/KennispuntTwente/microdata_catalogus.git ../microdata_catalogus
cd ../microdata_catalogus
uv sync
uv run python scripts/catalog.py update
cd ../microdata_synthThis creates ../microdata_catalogus/data/sqlite/catalogus.db and downloads PDF documentation locally. The generator reads the catalog's datasets, variables, join_keys, and datasets_fts tables, including variable metadata stored as JSON.
The initial update can take some time.
renv::restore()source("R/main.R")# Fuzzy search
search_datasets("polis")
# Explore catalog and keys
explore_catalogus()
explore_join_keys()
# Generate one dataset
secm <- generate_synthetic_data("SECMBUS", n_records = 500, seed = 42)
# Generate linked datasets
linked <- generate_linked_datasets(
dataset_names = c("SECMBUS", "SPOLISBUS"),
primary_dataset = "SECMBUS",
n_records = list(SECMBUS = 500, SPOLISBUS = 1000),
seed = 42
)
# Validate and export
check_referential_integrity(linked)
export_datasets(linked, output_dir = "output/sociaal_vangnet", format = "csv")generate_suite(
datasets = c("SECMBUS", "SPOLISBUS"),
primary_n = 500,
output_dir = "output/sociaal_vangnet",
format = c("csv", "rds"),
validate = TRUE,
verbose = TRUE
)R/
├── main.R # Entry point and public API
├── catalog_interface.R # SQLite catalog queries + metadata JSON parsing
├── data_generator.R # Core generation and constraint logic
├── validation.R # Type/range/integrity validation
└── export.R # CSV/RDS export helpers
output/ # Generated datasets
renv/ # Reproducible R environment
code_listvalues are used directly when availableusage_notesand descriptions are used for semantic detection (for example date fields)- explicit metadata
lengthoverrides parsed SQL type lengths
- linking is case-insensitive
- composite keys stay row-consistent when linking datasets
RINPERSOONSis generated from metadata code lists (one-letter source code), not numeric IDs
- dataset
periodis parsed and used as generation bounds - start/end pairs are enforced with synonym-aware matching:
- start patterns:
AANVANG,AANV,BEGIN,START,OPNAME - end patterns:
EINDE,EIND
- start patterns:
- only valid parseable date rows are adjusted (to avoid sentinel-code corruption)
- empty metadata
data_typevalues are handled gracefully num(N)andint(N)specifications are parsed correctly- referential integrity reporting works for composite person keys
Catalog exploration:
list_datasets()list_variables(dataset_name)search_datasets(pattern, mode = c("auto", "fuzzy", "fts"), limit = 10)search_datasets_fuzzy(pattern, limit = 10)search_datasets_fts(keyword, limit = 10)explore_catalogus()explore_join_keys()find_linkable_datasets(key_name = "RINPERSOON", key_type = NULL)find_datasets_by_key_type(key_type)get_dataset_metadata(dataset_name)get_dataset_keys(dataset_name)
Generation:
generate_synthetic_data(dataset_name, n_records = 100, seed = NULL, key_ids = NULL)generate_linked_datasets(dataset_names, n_records = 100, seed = NULL, primary_dataset = NULL, link_by = NULL)generate_quick(dataset_name = "GBAPERSOONTAB", n_records = 100, output_dir = "./output")generate_suite(datasets, primary_n = 100, output_dir = "./output", format = c("csv", "rds"), validate = TRUE, verbose = TRUE)
Validation and export:
validate_data(data, metadata, verbose = TRUE)check_referential_integrity(datasets, primary_dataset = NULL, key_cols = NULL, verbose = TRUE)check_duplicates(data, by = NULL)data_quality_report(data, metadata = NULL)export_data(data, output_dir, dataset_name = NULL, format = c("csv", "rds"), overwrite = TRUE)export_datasets(datasets, output_dir, format = c("csv", "rds"), organize_by_dataset = FALSE, verbose = TRUE)load_data(filepath, format = NULL)create_data_dictionary(data, metadata = NULL, output_file = NULL)summary_statistics(data, output_file = NULL)
- R 4.4.1+
microdata_catalogusSQLite database at../microdata_catalogus/data/sqlite/catalogus.db- Required packages (auto-checked in
R/main.R):dplyr,purrr,stringr,readrDBI,RSQLitecli,glue,tibble,rlang,jsonlite
- This repository generates synthetic data only; it does not connect to or contain restricted CBS microdata.
- Output quality depends on the catalog metadata and generation heuristics. Generated values and relationships can be unrealistic or internally inconsistent.
- Passing validation here does not establish correctness on real data. Review generated data and code before relying on them.
- The catalog is maintained separately; schema or metadata changes there may require updates here.
- Repository-specific Copilot feature and behavior documentation:
.github/copilot-instructions.md
- Repository Copilot skills:
.github/skills/synth-linking-integrity/SKILL.md.github/skills/synth-catalog-generation/SKILL.md.github/skills/synth-test-authoring/SKILL.md.github/skills/synth-docs-examples-sync/SKILL.md
MIT