themeparks-backfill: --since/--until, incremental reruns, final days only; changeRows() exposes opening - #55
Merged
Merged
Conversation
…rough
changeRows() flattened the raw history envelope to { entityId, row } and threw
away `opening`, the state each entity was in at the start of the range. Without
it the stretch before an entity's first change has no known status, which on a
night a ride runs past midnight is real operating time: Space Mountain on
2026-09-26 loses 63 seconds and no longer agrees with its own daily row.
The result is still the same async generator, so `for await` is unchanged. It
now also carries `opening`, keyed by entity id, readable once the response has
arrived (after iterating, or after `await changes.load()`), for one request.
HistorySpan gains `finalThrough`: the earlier of recordedTo and
retrievableThrough, the newest day whose daily rows will not change again.
The Space Mountain captures are shared byte for byte with the Python SDK.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…written Three defects from one customer-style run of 8.3.1, each reproduced against the live API first: 1. No date range. `--since` was an unknown option, so a key that reaches the whole archive downloaded all of it every time. `--since` and `--until` take a real calendar day as YYYY-MM-DD, both inclusive; since > until is refused before any request. A rerun accepts the same --since or a later one, and refuses one earlier than the day the file was asked to start from, or one that would leave a gap, saying what to do. 2. Reruns never updated. A finished park printed "already complete" and exited 0 without asking for a day. It now continues from the day after its last one; an interrupted run still resumes at its page boundary, and one interrupted before its first page records where it began. 3. Partial days, kept forever. A run ended at retrievableThrough, usually today. It now ends at span().finalThrough and says when it holds days back; the next run adds them once final. State moves to version 2: an 8.3 file has its last seven days before the old end removed (the rest byte for byte) and fetched again, and one wholly inside that window is downloaded again. Also: a finished state with a foreign column layout is refused instead of appended to; a state whose data file is gone starts again; and a continued file whose next day is older than the key may read is refused rather than jumping forward and leaving a gap. Eighteen new mutants cover these, and all 38 are killed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
README and CHANGELOG (Unreleased) for the backfill and history changes. The library example now ends at span().finalThrough, so it writes no partial days either. A test pins the README queue table's waitTime to the spec's `number`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
days() now awaits whatever onPage returns before requesting the next page, so a checkpoint written there is on disk before anything else can happen. The type is widened to `unknown`, so existing callbacks still compile. changeRows().opening is no longer enumerable, so spreading, Object.keys or a logger walking the result never trips the getter early. After a failed request it says the request failed rather than suggesting load(). The type docs explain `degraded`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…licates The state file was written only when a run ended or failed in a way it caught. A Ctrl-C, a SIGTERM or a kill during a multi-page download left the previous state on disk, so the next run fetched every page since and appended it a second time. - The state is written atomically (temp file, then rename) after every complete page, and once before the first request, so an extending run is marked unfinished before it asks for anything. - Each checkpoint waits for the rows before it to reach the file and records the file's size. A resume first cuts the file back to that size, so a page cut off mid-write is fetched again rather than appended twice. A file shorter than its checkpoint is refused. A state with no size that resumes on its last day drops that day first. - days() awaits onPage, so the checkpoint is on disk before the next request. - A lock beside the state file refuses a second run on the same output while one is running; a lock left by a process that is gone is taken over. - SIGINT and SIGTERM release the locks and exit (130, 143). Nothing else needs flushing: the last checkpoint is already on disk. Tested by abandoning a run mid-page, as a kill would, and resuming it. The stream-error test now reaches the stream (the lock is taken first), and five stale mutants are updated alongside four new ones. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wording - The checkpoint field is `size`, as in Python. A rerun cuts the file back to it and says "discarding the last N bytes of X: written after the last checkpoint, and fetched again now". A file shorter than its checkpoint is refused in the same words. - A state with no size is cut back by date: a finished file to its end, an unfinished one to its page boundary, or a whole page (31 days) before its last day when it has none. Resuming at the last day, as before, lost the days of entities that had not reached it. - `start` is the first day the file covers, after the key's floor; `since` is the start it was asked for. A --since before the first day is refused unless it is that asked-for start, with Python's wording. - The lock and anonymous-access messages match Python's, and the anonymous notice says only final days are written, usually 4 or 5 per park. - An 8.3 file being upgraded says what is done to it, and an --until the file already holds no longer blames the key's plan. Tests for the two mutants an independent sweep found surviving (the cut boundary, and a run that adds no rows keeping lastDay); 45/45 mutants killed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…g safely A day through recordedTo is one the archive has recorded and is fetched once; the archive can occasionally re-record a past day after a feed repair, so the README now says how to fetch a range again. Documents that stopping a run is safe at any point, what `opening.degraded` means, and, in the CHANGELOG, that this is 8.4.0 with `HistorySpan` gaining a required field. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
A customer-style run of
themeparks-backfill8.3.1 found three problems in the command and one in the history helpers. Each was reproduced against the live API before it was fixed. The Python SDK ships the same fixes (ThemeParks/ThemeParks_Python#33), and the two behave the same except where noted below.--sincewasUnknown option, so you always got everything your key reaches. New flags:--since YYYY-MM-DDand--until YYYY-MM-DD. Both days are inclusive, each must be a real calendar day (2025-02-30,2025-1-1and20250101are rejected), and--sinceafter--untilis refused before any request is made.already completeand exited 0 without fetching anything, so a nightly cron looked healthy but never updated. Now a finished file continues from the day after its last one. An interrupted run still resumes at its page boundary.retrievableThrough, usually today. Today's row is the day so far, and the API docs say recording runs 2 to 3 days behind live data, so the last few days of every file were written while they could still change. Now:span().finalThrough(the earlier ofrecordedToandretrievableThrough) and says when it held days back;history.changeRows()threw awayopening, the state in force at the start of the range. It is now exposed without breaking anything: the result is the same async generator yielding the same{ entityId, row }entries, and it also carriesopening, keyed by entity id, readable after iterating or afterawait changes.load(), for one request. (history.changes()returns the raw envelope and always had it.)waitTime: checked, no mismatch here. The README and the generated types both saynumber, matching the spec. A test now pins the README table to that.Also fixed along the way (each has a test and a committed mutant)
discarding the last N bytes of X: written after the last checkpoint, and fetched again nowand cuts the file back to it. An extension is marked unfinished before it asks for anything. SIGTERM exits 143 (Ctrl-C 130).days()now awaitsonPage, so the checkpoint is on disk before the next request.Compatibility
HistorySpangains a required field,finalThrough, so code that builds one by hand needs to add it.HistoryChangesandHistoryOpeningare exported types.onPagemay now return a promise.startis the first day the file covers (after the key's window),sincethe start it was asked for,endthe newest recorded day, andsizethe file's size at the checkpoint.end(kept rows are copied byte for byte) and fetches those days again, and says so. A file wholly inside that window, as every anonymous run's is, is downloaded again.--sinceon a rerun: the one the file was asked to start from, or one inside the file, is accepted (a fixed or rolling cron line works); any other before the file's first day, or one that would leave a gap, is refused with what to do.recordedTois one the archive has recorded and is fetched once. The archive can re-record a past day after a feed repair; the README says how to fetch a range again.The Python library's matching PR follows the same rules and messages.
Verification
npm test: 300 passed.npm run lint,npm run typecheck,prettier --check .,npm run buildandnpm run test:packageare clean.test/mutation/run.mjs: 45/45 killed.(entityId, date));recordedTo; running it again reportedup to datewith no history request;--since/--untilwrote only those days; an earlier--sincewas refused with the file untouched;--sinceafter--untilexited 2 before any request;changeRows().openingfor Space Mountain on 2026-09-26 isOPERATINGat midnight, as in the capture.🤖 Generated with Claude Code