feat: compare databases, mirror volumes and seed profiles - #41
Merged
Merged
Conversation
✅ PR title follows the required formatCurrent title: |
Compare and mirror
- compare: rows, size, delta and column drift per table between any two
connections, engines may differ; --counts estimate falls back to exact
counts for tables without statistics
- mirror: seed a target to follow a source's volumes (--scale, --max-rows,
--tables, topup or reset) with a --dry-run plan and sample rows; refuses
when source and target are the same database
Seed profiles
- value rules by column pattern or per column: template ({{auto}}, {{seq}},
{{run}}, any generator), faker, value, oneOf, setNull
- saved in ~/.config/seedstorm/profiles.yaml and used via --profile on seed,
gaps, generate and mirror, the profile command, the TUI and the web builder
- rules skip keys and incompatible types; single and multi-column UNIQUE
constraints stay satisfied
Seeding engine
- one engine behind CLI, TUI and web: seeder.Seed (seed, gaps, generate)
and seeder.Fill (mirror: refused rows regenerated, stuck tables reported)
- rows are generated and written 20k at a time, through COPY on Postgres;
PK pools are sampled and key sets bounded, so memory stays flat
(7.2M-row mirror: 147MB peak, 600k-row seed: 97MB)
- re-seeding appends past existing ids, UNIQUE values and composite keys;
Postgres sequences are advanced after every run
- default batch size 1000, split to stay under 65,535 placeholders and ~1MB
Generate and export
- streamed through internal/dataio (yaml, json, sql, csv) with the same file
layout; 300k rows use ~100MB instead of 1.6-2.1GB
- SQL output carries literal values escaped per engine and loads as-is
Web UI
- Compare and Profiles pages, profile picker in the workspace, nav drawer
on narrow screens
- static assets revalidate with ETags; job output capped at 20MB, request
bodies limited, idle session connections closed
Fixes
- integer *_time/*_date columns got time strings; string PKs ignored varchar(n)
- MySQL clone-schema lost index prefix lengths; MySQL 5.7 rejected a cloned
nullable TIMESTAMP
- profile table names missed MySQL identifier case; value: true failed on
TINYINT(1)
Tests: binary-driven evals on Postgres 13/15/17 and MySQL 5.7/8.0/8.4 covering
Keycloak's 87-table schema, mirror, sequences, memory limits at 300k rows and
SQL round trips; integration timeout raised to 900s.
machado144
force-pushed
the
feat/compare-mirror-and-seed-rules
branch
from
September 14, 2026 22:03
50154c5 to
830e553
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
compare: row counts, size, delta and column drift per table between any two connections. Read-only on both sides; tables match across engines ignoring case.mirror: seeds a target so its volumes follow a source (--scale,--max-rows,--tables,topuporreset).--dry-runshows the plan and sample rows. Refuses when source and target are the same database.template({{auto}},{{seq}},{{run}}, any generator),faker,value,oneOfandsetNull. Profiles are saved to~/.config/seedstorm/profiles.yamland used via--profileon seed/gaps/generate/mirror, in the TUI, and in the web UI.internal/seeder): tables fill in chunks and rejected rows are regenerated. A table that can't make progress is reported aspartial/failedand the run continues.seeder.Seedengine for CLI, TUI and web); Postgres chunks go throughCOPY; key state is read once per table with sampled PK pools and Bloom-filter key sets.seedof 600k rows on Postgres: 41.9s / 585MB → 7.3s / 97MBinternal/dataio): files are written and read a chunk at a time, same YAML/JSON layout as before. 300k rows: generate 1.65GB → 99MB, export 2.14GB → 80MB. CSV added to generate, YAML/JSON input to export.--out), job requests are size-limited, and idle session connections close after 2 minutes.Fixes found while testing against Keycloak and the DB version matrix:
*_time/*_datecolumns stored as integers got time stringsvarchar(n)TIMESTAMPUNIQUE (realm_id, username)) were ignored; tuples are now kept distinct and rows that can't be made distinct are dropped with a warningvalue: truefailed on MySQLTINYINT(1)--counts estimatereported unknown or stale-zero counts on never-analyzed tables (pg13); those tables are now counted exactlygenerate --format sql,export --format sqland dry runs had$1/?placeholders and no values; it now writes literals escaped per engine and loads as-isCI integration timeout is raised to 900s; the suite now takes about 5 minutes.
Test plan
go test -race ./...) for rules, compare/plan, seeder retry decisions (scripted driver), profiles store, faker existing-row handling, TUI mirror model, and web APIspostgres:17container, and mysql8.0 → freshmysql:8.4container; 1× and 2× exact on all 87 tables, 0 custom-field violations in the target;generate --profilechecked on bothseed/gaps --fillof 300k rows under 150MB peak (old code: 299MB) and an 80-column table at the default batch size (fails on MySQL without placeholder splitting)generate/exportof 300k rows under 150MB; exported and generated SQL loads on both engines with quotes, backslashes, newlines, unicode and NULL intact (drop the MySQL backslash escaping and it fails)