Skip to content

feat: compare databases, mirror volumes and seed profiles - #41

Merged
machado144 merged 1 commit into
mainfrom
feat/compare-mirror-and-seed-rules
Sep 15, 2026
Merged

machado144 merged 1 commit into
mainfrom
feat/compare-mirror-and-seed-rules

Conversation

@machado144

@machado144 machado144 commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • compare: row counts, size, delta and column drift per table between any two connections. Read-only on both sides; tables match across engines ignoring case.
  • mirror: seeds a target so its volumes follow a source (--scale, --max-rows, --tables, topup or reset). --dry-run shows the plan and sample rows. Refuses when source and target are the same database.
  • Seed profiles: value rules applied by column pattern or per column. Actions are template ({{auto}}, {{seq}}, {{run}}, any generator), faker, value, oneOf and setNull. Profiles are saved to ~/.config/seedstorm/profiles.yaml and used via --profile on seed/gaps/generate/mirror, in the TUI, and in the web UI.
  • Resilient inserts (internal/seeder): tables fill in chunks and rejected rows are regenerated. A table that can't make progress is reported as partial/failed and the run continues.
  • Large volumes with flat memory: seed, gaps and mirror generate and write 20k rows at a time (one shared seeder.Seed engine for CLI, TUI and web); Postgres chunks go through COPY; key state is read once per table with sampled PK pools and Bloom-filter key sets.
    • 7.2M-row mirror: 3m23s / 1.48GB peak → 2m27s / 147MB
    • seed of 600k rows on Postgres: 41.9s / 585MB → 7.3s / 97MB
  • Streaming generate/export (internal/dataio): files are written and read a chunk at a time, same YAML/JSON layout as before. 300k rows: generate 1.65GB → 99MB, export 2.14GB → 80MB. CSV added to generate, YAML/JSON input to export.
  • Bounded serve: job output to the browser caps at 20MB (dry-run SQL notes what it left out; generate/export refuse and point at --out), job requests are size-limited, and idle session connections close after 2 minutes.
  • Faster MySQL inserts: default batch size 100 → 1000 (200k rows: 36s → 9s); batches split to stay under 65,535 placeholders and ~1MB.
  • Postgres sequences behind SERIAL/IDENTITY columns are advanced after seed, gaps and mirror, so the application's own inserts don't collide with seeded ids.
  • Web: new Compare and Profiles pages, a profile select in the workspace, and a nav drawer on narrow screens.
  • Re-seeding a populated DB now appends: ids, UNIQUE sequences and composite keys continue past existing rows.

Fixes found while testing against Keycloak and the DB version matrix:

  • *_time/*_date columns stored as integers got time strings
  • string PKs ignored varchar(n)
  • MySQL clone-schema lost index prefix lengths
  • MySQL 5.7 rejected a cloned nullable TIMESTAMP
  • multi-column UNIQUE constraints (Keycloak's UNIQUE (realm_id, username)) were ignored; tuples are now kept distinct and rows that can't be made distinct are dropped with a warning
  • explicit profile table/column names didn't match MySQL's upper-case identifiers; value: true failed on MySQL TINYINT(1)
  • --counts estimate reported unknown or stale-zero counts on never-analyzed tables (pg13); those tables are now counted exactly
  • serve re-downloaded ~870KB of scripts on every page; static assets now revalidate with ETags
  • SQL from generate --format sql, export --format sql and dry runs had $1/? placeholders and no values; it now writes literals escaped per engine and loads as-is

CI integration timeout is raised to 900s; the suite now takes about 5 minutes.

seedstorm compare --source-dsn "$PROD" --target-dsn "$STAGE" --only-diff
seedstorm mirror  --source-dsn "$PROD" --target-dsn "$STAGE" --scale 5 --max-rows 2000000 --profile loadtest --dry-run
name: loadtest
rules:
  - column: "*email*"
    template: "lt+{{seq}}.{{run}}@example.test"
tables:
  users:
    columns:
      role: { value: guest }

Test plan

  • Unit tests (go test -race ./...) for rules, compare/plan, seeder retry decisions (scripted driver), profiles store, faker existing-row handling, TUI mirror model, and web APIs
  • Binary-driven integration evals on Postgres + MySQL: re-seed, compare/mirror end to end, seeder rejections, Keycloak 87-table schema (pg, mysql, mysql→pg), web compare/mirror/profiles
  • Full integration suite green (282 subtests); new evals also green on pg13 + mysql5.7 and pg17 + mysql8.4
  • Real replication with every rule kind: Keycloak on pg15 → fresh postgres:17 container, and mysql8.0 → fresh mysql:8.4 container; 1× and 2× exact on all 87 tables, 0 custom-field violations in the target; generate --profile checked on both
  • Compare and Profiles pages driven in Playwright; zero horizontal overflow at 320–1920px
  • Evals: seed/gaps --fill of 300k rows under 150MB peak (old code: 299MB) and an 80-column table at the default batch size (fails on MySQL without placeholder splitting)
  • Evals: generate/export of 300k rows under 150MB; exported and generated SQL loads on both engines with quotes, backslashes, newlines, unicode and NULL intact (drop the MySQL backslash escaping and it fails)
  • Web generate → Export this → export SQL, dry runs and the 20MB refusal driven in Playwright
  • 7.2M-row mirror benchmark (accounts → orders → order_items): no orphans, unique emails, CHECKs hold, sequences advanced
  • golangci-lint, gofumpt, gauntlet (structlint, dupehound on changed lines)

@github-actions

Copy link
Copy Markdown

✅ PR title follows the required format

Current title: feat: compare databases, mirror volumes and seed profiles

Compare and mirror
- compare: rows, size, delta and column drift per table between any two
  connections, engines may differ; --counts estimate falls back to exact
  counts for tables without statistics
- mirror: seed a target to follow a source's volumes (--scale, --max-rows,
  --tables, topup or reset) with a --dry-run plan and sample rows; refuses
  when source and target are the same database

Seed profiles
- value rules by column pattern or per column: template ({{auto}}, {{seq}},
  {{run}}, any generator), faker, value, oneOf, setNull
- saved in ~/.config/seedstorm/profiles.yaml and used via --profile on seed,
  gaps, generate and mirror, the profile command, the TUI and the web builder
- rules skip keys and incompatible types; single and multi-column UNIQUE
  constraints stay satisfied

Seeding engine
- one engine behind CLI, TUI and web: seeder.Seed (seed, gaps, generate)
  and seeder.Fill (mirror: refused rows regenerated, stuck tables reported)
- rows are generated and written 20k at a time, through COPY on Postgres;
  PK pools are sampled and key sets bounded, so memory stays flat
  (7.2M-row mirror: 147MB peak, 600k-row seed: 97MB)
- re-seeding appends past existing ids, UNIQUE values and composite keys;
  Postgres sequences are advanced after every run
- default batch size 1000, split to stay under 65,535 placeholders and ~1MB

Generate and export
- streamed through internal/dataio (yaml, json, sql, csv) with the same file
  layout; 300k rows use ~100MB instead of 1.6-2.1GB
- SQL output carries literal values escaped per engine and loads as-is

Web UI
- Compare and Profiles pages, profile picker in the workspace, nav drawer
  on narrow screens
- static assets revalidate with ETags; job output capped at 20MB, request
  bodies limited, idle session connections closed

Fixes
- integer *_time/*_date columns got time strings; string PKs ignored varchar(n)
- MySQL clone-schema lost index prefix lengths; MySQL 5.7 rejected a cloned
  nullable TIMESTAMP
- profile table names missed MySQL identifier case; value: true failed on
  TINYINT(1)

Tests: binary-driven evals on Postgres 13/15/17 and MySQL 5.7/8.0/8.4 covering
Keycloak's 87-table schema, mirror, sequences, memory limits at 300k rows and
SQL round trips; integration timeout raised to 900s.
@machado144
machado144 force-pushed the feat/compare-mirror-and-seed-rules branch from 50154c5 to 830e553 Compare September 14, 2026 22:03
@machado144
machado144 merged commit 6e98719 into main Sep 15, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant