Skip to content

themeparks-backfill: --since/--until, incremental reruns, final days only; changeRows() exposes opening - #55

Merged
cubehouse merged 7 commits into
mainfrom
fix/backfill-incremental
Sep 29, 2026
Merged

cubehouse merged 7 commits into
mainfrom
fix/backfill-incremental

Conversation

@cubehouse

@cubehouse cubehouse commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

What this fixes

A customer-style run of themeparks-backfill 8.3.1 found three problems in the command and one in the history helpers. Each was reproduced against the live API before it was fixed. The Python SDK ships the same fixes (ThemeParks/ThemeParks_Python#33), and the two behave the same except where noted below.

  1. No date range. --since was Unknown option, so you always got everything your key reaches. New flags: --since YYYY-MM-DD and --until YYYY-MM-DD. Both days are inclusive, each must be a real calendar day (2025-02-30, 2025-1-1 and 20250101 are rejected), and --since after --until is refused before any request is made.
  2. Reruns never updated. Once a park was complete, running it again printed already complete and exited 0 without fetching anything, so a nightly cron looked healthy but never updated. Now a finished file continues from the day after its last one. An interrupted run still resumes at its page boundary.
  3. The last days were partial, and stayed partial. Runs ended on retrievableThrough, usually today. Today's row is the day so far, and the API docs say recording runs 2 to 3 days behind live data, so the last few days of every file were written while they could still change. Now:
    • a run ends at the new span().finalThrough (the earlier of recordedTo and retrievableThrough) and says when it held days back;
    • the next run adds those days once they are final;
    • each day is fetched once, as the archive recorded it, so the file only ever gets appended to.
  4. history.changeRows() threw away opening, the state in force at the start of the range. It is now exposed without breaking anything: the result is the same async generator yielding the same { entityId, row } entries, and it also carries opening, keyed by entity id, readable after iterating or after await changes.load(), for one request. (history.changes() returns the raw envelope and always had it.)
  5. waitTime: checked, no mismatch here. The README and the generated types both say number, matching the spec. A test now pins the README table to that.

Also fixed along the way (each has a test and a committed mutant)

  • An interrupted run appended days twice. The state was written only at the end of a run or on an error the command caught, so Ctrl-C, SIGTERM or a kill during a multi-page download or a nightly extension left an older state, and the rerun appended the same days again. The state is now written atomically (temp file, then rename) before the first request and after every page, with the data file's size; a rerun first prints discarding the last N bytes of X: written after the last checkpoint, and fetched again now and cuts the file back to it. An extension is marked unfinished before it asks for anything. SIGTERM exits 143 (Ctrl-C 130). days() now awaits onPage, so the checkpoint is on disk before the next request.
  • Two runs on one output at once are refused with a lock beside the state file ("another themeparks-backfill is writing this park"); a lock left by a process that is gone is taken over.
  • A finished state file with a different column layout was appended to: the whole archive a second time, under a second header, exit 0. It is now refused, like an unfinished one.
  • If the data file had been deleted but its state file remained, the next run carried on part-way through. It now downloads the park again.
  • A continued file whose next day is older than the key may read is refused rather than carried forward past a gap.

Compatibility

  • To be released as 8.4.0. HistorySpan gains a required field, finalThrough, so code that builds one by hand needs to add it. HistoryChanges and HistoryOpening are exported types. onPage may now return a promise.
  • State files move to version 2. start is the first day the file covers (after the key's window), since the start it was asked for, end the newest recorded day, and size the file's size at the checkpoint.
  • Version 1 files from this SDK are upgraded, not refused. Nobody recorded which of their newest days were final, so the upgrade removes rows dated within 7 days of the old end (kept rows are copied byte for byte) and fetches those days again, and says so. A file wholly inside that window, as every anonymous run's is, is downloaded again.
  • --since on a rerun: the one the file was asked to start from, or one inside the file, is accepted (a fixed or rolling cron line works); any other before the file's first day, or one that would leave a gap, is refused with what to do.
  • A day through recordedTo is one the archive has recorded and is fetched once. The archive can re-record a past day after a feed repair; the README says how to fetch a range again.

The Python library's matching PR follows the same rules and messages.

Verification

  • npm test: 300 passed. npm run lint, npm run typecheck, prettier --check ., npm run build and npm run test:package are clean.
  • test/mutation/run.mjs: 45/45 killed.
  • Interruption is tested by abandoning a run mid-page, as a kill would, and resuming it (NDJSON and CSV, first run and extension).
  • Live, against Bellewaerde (one small park) and Space Mountain:
    • a keyed four-page run was stopped with SIGINT after its first page: exit 130, state at the page boundary, lock released; the rerun resumed there and the file came out byte for byte identical to an uninterrupted run (3,270 rows, no duplicated (entityId, date));
    • an 8.3.1 anonymous file was downloaded again and stopped at recordedTo; running it again reported up to date with no history request;
    • --since/--until wrote only those days; an earlier --since was refused with the file untouched; --since after --until exited 2 before any request;
    • changeRows().opening for Space Mountain on 2026-09-26 is OPERATING at midnight, as in the capture.

🤖 Generated with Claude Code

cubehouse and others added 7 commits September 28, 2026 21:46
…rough

changeRows() flattened the raw history envelope to { entityId, row } and threw
away `opening`, the state each entity was in at the start of the range. Without
it the stretch before an entity's first change has no known status, which on a
night a ride runs past midnight is real operating time: Space Mountain on
2026-09-26 loses 63 seconds and no longer agrees with its own daily row.

The result is still the same async generator, so `for await` is unchanged. It
now also carries `opening`, keyed by entity id, readable once the response has
arrived (after iterating, or after `await changes.load()`), for one request.

HistorySpan gains `finalThrough`: the earlier of recordedTo and
retrievableThrough, the newest day whose daily rows will not change again.

The Space Mountain captures are shared byte for byte with the Python SDK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…written

Three defects from one customer-style run of 8.3.1, each reproduced against
the live API first:

1. No date range. `--since` was an unknown option, so a key that reaches the
   whole archive downloaded all of it every time. `--since` and `--until` take
   a real calendar day as YYYY-MM-DD, both inclusive; since > until is refused
   before any request. A rerun accepts the same --since or a later one, and
   refuses one earlier than the day the file was asked to start from, or one
   that would leave a gap, saying what to do.

2. Reruns never updated. A finished park printed "already complete" and exited
   0 without asking for a day. It now continues from the day after its last
   one; an interrupted run still resumes at its page boundary, and one
   interrupted before its first page records where it began.

3. Partial days, kept forever. A run ended at retrievableThrough, usually
   today. It now ends at span().finalThrough and says when it holds days back;
   the next run adds them once final. State moves to version 2: an 8.3 file has
   its last seven days before the old end removed (the rest byte for byte) and
   fetched again, and one wholly inside that window is downloaded again.

Also: a finished state with a foreign column layout is refused instead of
appended to; a state whose data file is gone starts again; and a continued
file whose next day is older than the key may read is refused rather than
jumping forward and leaving a gap.

Eighteen new mutants cover these, and all 38 are killed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
README and CHANGELOG (Unreleased) for the backfill and history changes. The
library example now ends at span().finalThrough, so it writes no partial days
either. A test pins the README queue table's waitTime to the spec's `number`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
days() now awaits whatever onPage returns before requesting the next page, so
a checkpoint written there is on disk before anything else can happen. The
type is widened to `unknown`, so existing callbacks still compile.

changeRows().opening is no longer enumerable, so spreading, Object.keys or a
logger walking the result never trips the getter early. After a failed request
it says the request failed rather than suggesting load(). The type docs explain
`degraded`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…licates

The state file was written only when a run ended or failed in a way it caught.
A Ctrl-C, a SIGTERM or a kill during a multi-page download left the previous
state on disk, so the next run fetched every page since and appended it a
second time.

- The state is written atomically (temp file, then rename) after every
  complete page, and once before the first request, so an extending run is
  marked unfinished before it asks for anything.
- Each checkpoint waits for the rows before it to reach the file and records
  the file's size. A resume first cuts the file back to that size, so a page
  cut off mid-write is fetched again rather than appended twice. A file
  shorter than its checkpoint is refused. A state with no size that resumes on
  its last day drops that day first.
- days() awaits onPage, so the checkpoint is on disk before the next request.
- A lock beside the state file refuses a second run on the same output while
  one is running; a lock left by a process that is gone is taken over.
- SIGINT and SIGTERM release the locks and exit (130, 143). Nothing else needs
  flushing: the last checkpoint is already on disk.

Tested by abandoning a run mid-page, as a kill would, and resuming it. The
stream-error test now reaches the stream (the lock is taken first), and five
stale mutants are updated alongside four new ones.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wording

- The checkpoint field is `size`, as in Python. A rerun cuts the file back to
  it and says "discarding the last N bytes of X: written after the last
  checkpoint, and fetched again now". A file shorter than its checkpoint is
  refused in the same words.
- A state with no size is cut back by date: a finished file to its end, an
  unfinished one to its page boundary, or a whole page (31 days) before its
  last day when it has none. Resuming at the last day, as before, lost the days
  of entities that had not reached it.
- `start` is the first day the file covers, after the key's floor; `since` is
  the start it was asked for. A --since before the first day is refused unless
  it is that asked-for start, with Python's wording.
- The lock and anonymous-access messages match Python's, and the anonymous
  notice says only final days are written, usually 4 or 5 per park.
- An 8.3 file being upgraded says what is done to it, and an --until the file
  already holds no longer blames the key's plan.

Tests for the two mutants an independent sweep found surviving (the cut
boundary, and a run that adds no rows keeping lastDay); 45/45 mutants killed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…g safely

A day through recordedTo is one the archive has recorded and is fetched once;
the archive can occasionally re-record a past day after a feed repair, so the
README now says how to fetch a range again. Documents that stopping a run is
safe at any point, what `opening.degraded` means, and, in the CHANGELOG, that
this is 8.4.0 with `HistorySpan` gaining a required field.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cubehouse
cubehouse merged commit 9e55569 into main Sep 29, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant