feat(backfill): themeparks-backfill, and the defects porting it exposed - #51
Merged
Merged
Conversation
…d in it The Python SDK shipped this command first. Porting it, diffing the two outputs over the same park, and then reviewing the result from six angles found around twenty defects between them. Two independent implementations reading one API disagree in exactly the places one of them is wrong. Magic Kingdom's full five-year archive now comes back byte for byte identical from both SDKs: 94,223 rows, 41 columns, the only differences being today's row, which grows as the day elapses. npm install themeparks npx themeparks-backfill "magic kingdom" A park or a destination, by name or id, one file per park. NDJSON by default, --format csv for one wide row per entity per day. Every row carries parkId, parkName, entityId, entityName and entityType, so two files load into one table and (entityId, date) is the natural key. The entity name is the one the history response gave for those rows, not the park's current children list, because rides get renamed and today's name on a row from three years ago rewrites the record. Files are named for the park's id, because names change. --list needs no key, so you can find your park before deciding whether to pay. Resumable: it checkpoints against the hourly history budget and exits 75, so a timer retries rather than alerting. The checkpoint is the day the server's own `next` URL starts on, never the newest row written -- an entity that stopped reporting has no rows for the tail days of its page, so a row-derived checkpoint re-fetches days already in the file. That boundary is reported through a new `onPage` hook on days(), since the page boundary is the server's answer to "where do I carry on" and the rows cannot tell you. The state file is <parkId>.<format>.backfill-state.json and records the SDK, its version, a state version and a fingerprint of the exact header. Anything that does not match is refused with a message saying why. One state file for two formats doubled every row on an ndjson -> csv -> ndjson round trip; no record of the header let a 19-column file resume under a 41-column build; and the two SDKs' keys differed only on the interrupted path, so the safe paths interoperated and a cross-SDK resume produced 172 rows where 108 belonged. Three things only a review caught. A failed write was reported as success: Node hands `end`'s callback the stream's error and the callback took no arguments, so ENOSPC mid-download printed "done", recorded complete, and exited 0 with a truncated file. A failure on a resumed run deleted every row already downloaded, because `written === 0` means this process wrote nothing rather than the file is empty. And the executable's entry-point guard compared import.meta.url against process.argv[1], which npm makes a symlink, so `npx themeparks-backfill --version` printed nothing and exited 0 -- green on all 186 unit tests and working in the repo. The runner is its own file now and scripts/check-package.ts packs, installs and runs the binary in CI. The vendored OpenAPI schema was stale, so unknownMinutes, inParkHours and extremeWaits were on every row the API returns and in none of the types. The CSV column list is generated from the spec, the nightly drift job commits it alongside the schema, and test/fixtures/csv_contract.json is asserted by both SDKs so the two headers cannot diverge again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
cubehouse
force-pushed
the
feat/backfill-parity
branch
from
September 28, 2026 16:18
511efbd to
c70ed00
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
themeparks-backfillfor the JavaScript SDK. The Python SDK shipped it first; this is the same tool, and the two now write byte-for-byte identical CSVs. Magic Kingdom's full five-year archive: 94,223 rows, 41 columns, identical from both, the only differences being today's row, which grows as the day elapses.npm install themeparks npx themeparks-backfill "magic kingdom"Nothing published has ever carried a
bin, so there is no compatibility burden here and 8.3.0 stays a minor.What six reviews found
A failed write was reported as success. Node hands
end's callback the stream's error; the callback took no arguments and resolved regardless, so on ENOSPC mid-download the command printeddone: N rows, recordedcomplete: trueand exited 0 with a truncated file no rerun would continue. The stream also had no'error'listener until the flush, so an earlier failure became an unhandled'error'event that killed the run and left five of a destination's six parks unattempted.A failure on a resumed run deleted every row already downloaded, because
written === 0means "this process wrote nothing", not "the file is empty".The entry-point guard was wrong for an installed binary. It compared
import.meta.urlagainstprocess.argv[1], and npm installs abinas a symlink — sonpx themeparks-backfill --versionprinted nothing and exited 0. Green on all 186 unit tests and working when run inside the repo. Only packing and installing the tarball showed it. The runner is its own file now, andscripts/check-package.tspacks, installs and runs the binary in CI.The vendored OpenAPI schema was stale.
unknownMinutes,inParkHoursandextremeWaitsare on every row the API returns and were in none of the types. The column list is generated from the spec now, andspec-drift.ymlcommits it alongside the schema — it previously committed only the schema, which would have silently restored the exact defect this release fixes, with nothing red anywhere.HistoryPagewas never exported although the changelog advertised it (dist/index.d.tshad zero references).Plus: a bare
\rin an entity name written unquoted, so one row parsed as two;--listwith no value exiting 2 though the help says[TEXT]; no-h;東京listing all 127 parks becausenormalizefolds it to''; an empty--api-keycounting as a key; no-arguments fetching/destinationsbefore saying so, which exits 75 with no network;--versionprinting a bare number; the 404 hint sitting where nothing could reach it; and a unique substring being refused while the Python SDK downloaded it.The shared contract
test/fixtures/csv_contract.jsonis asserted by both SDKs. Before it existed this SDK wrote 32 columns and Python wrote 41, withinParkScheduledMinutesagainstinParkHoursScheduledMinutes, and the test that claimed to check it compared four strings. Both now compute the same fingerprint,1b6ea478049dc7ca, over the same header.The state file is
<parkId>.<format>.backfill-state.jsoncarrying the SDK, its version, a state version and that fingerprint, refusing anything that does not match — including a state file written by the Python SDK, whose keys differ only on the interrupted path, which is how a cross-SDK resume produced 172 rows where 108 belonged.Tests
211 (from 186), none needing the network, and 13 of 13 targeted mutations killed — including a renamed column in the generated file, caught by the shared contract. Fixtures are real captures: a real nested 403 body, 101 real destinations, and two real consecutive pages of one request where the newest row, the page's last day and the resume point are three different dates.
One honest gap:
history_rate_limited.jsonis hand-written, and the SDK reads only theRetry-Afterheader, so the fixture's body earns nothing. The README now says so — the hourly-budget header contract is unverified against production, and capturing a real 429 costs an hour of a metered key's budget.Companion: ThemeParks_Python#29.
🤖 Generated with Claude Code