feat: parallel seeding with live progress, access checks, counts snapshots and object clone - #42
Merged
Merged
Conversation
…shots and object clone - Seed/gaps/mirror write on --workers connections (default 4), FK-gated per table; runs report rows, rate and ETA (CLI every 2s, web live) - Web: connection count + privilege badge, missing-privilege warnings, search zoom-to-matches, navigator minimap, one-row action bar - snapshot command, --source-snapshot on compare/mirror; web export/import of counts and last comparison kept per pair - clone-schema --views/--routines/--triggers (--objects all) - Profiles: ignore globs honoured by every run, editor + workspace tab, profile file download/import
✅ PR title follows the required formatCurrent title: |
…13 access test - Guard the job log writer: concurrent seed writers panicked serve mid-run - Pack dependency levels so 150 tables open readable (zoom 0.15 → 0.40); spread search matches jump to the best one, with a match bar and an "Only matches" view; cap fit zoom for tiny schemas - Mobile workspace: rail no longer hidden under the canvas, no sideways scroll from wide previews - Postgres 13 grants CREATE on public to PUBLIC: access test sets its state - 150-table cross-referenced schema eval with 8 writers on both engines - Docs: new evals, safe local version matrix, layout and search behaviour
- Chunks and write queue bounded by measured row memory; key pools freed once no remaining table reads them (6KB rows 377MB -> 135MB; 100 tables 262MB -> 82MB) - --gen-workers / Generators: tables generate on separate cores with private random sources, respecting FK and near-cycle order (446k -> 1.79M rows/s) - --seed reproduces byte-identical data (table order, dates, sampling, enums) - Faster single-core generation (cached faker specs, O(1) self-ref parents) - Fix web job races (subscriber map, status reads) and progress event flood - Playwright e2e journeys with data-testid selectors; web job, memory, reproducibility and parallel-generation integration tests
- Queued rows cost at least half an average chunk row (4 writers: 157MB -> 117MB) - Deadlock / lock-timeout writes retry once - Memory tests read the child's VmHWM instead of rusage - CI job timeouts; benchmarks and do/don't lessons documented
- Cancel superseded PR runs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Feedback from seeding a 128-table staging DB: silent 9-minute runs, no idea what the user may do, unreadable graph, counts stuck in the tool.
Scale & progress
--workers(default 4; web: Tuning → Writers, Compare → Advanced): writes go to a pool while generation streams. A table writes only after every FK parent in the run finished; self-referencing tables stay ordered.1= old sequential path.--gen-workers(web: Tuning → Generators): tables generate on separate cores, each with its own random source. 32 tables in memory: 446k rows/s (1) → 1.79M rows/s (8). Off with--seed.--seedis now truly reproducible (sorted graph, seeded dates/pools/enum pick).Progress 120.0k/1.5M rows (7.8%) · 42.1k rows/s · ETA 33s, web bar with rows/rate/ETA (workspace and compare).Numbers and how to reproduce them:
docs/benchmarks.md.Web UI
full access/limited/read-only) fromhas_table_privilege/SHOW GRANTS; workspace lists missing privileges, marks no-INSERT tables, warns per mode and for the clone/mirror target.Counts, clone, profiles
seedstorm snapshot+--source-snapshoton compare/mirror; Compare exports either side (JSON/YAML, download/copy) and imports file/drop/paste.clone-schema --views --routines --triggers(--objects all), also in the workspace clone mode.ignore:globs honoured by seed/gaps/generate/mirror (empty required parent → clear error), editor on Profiles, Ignored tab in the workspace; profile YAML download/file import.Tests & CI
-race): Postgres 13/15/17 × MySQL 5.7/8.0/8.4 — concurrent vs sequential seed, 150-table cross-referenced schema, memory bounds (child VmHWM), access with limited users, snapshot compare/mirror, object clone, web jobs.serve(e2e/,make test-e2e), selectors bydata-testid.timeout-minuteson every job, superseded PR runs cancelled, e2e Node pinned (e2e/.nvmrc) with faster registry retries and a cached Chromium.