Skip to content

feat(backfill): themeparks-backfill, and the defects porting it exposed - #51

Merged
cubehouse merged 1 commit into
mainfrom
feat/backfill-parity
Sep 28, 2026
Merged

cubehouse merged 1 commit into
mainfrom
feat/backfill-parity

Conversation

@cubehouse

@cubehouse cubehouse commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

themeparks-backfill for the JavaScript SDK. The Python SDK shipped it first; this is the same tool, and the two now write byte-for-byte identical CSVs. Magic Kingdom's full five-year archive: 94,223 rows, 41 columns, identical from both, the only differences being today's row, which grows as the day elapses.

npm install themeparks
npx themeparks-backfill "magic kingdom"

Nothing published has ever carried a bin, so there is no compatibility burden here and 8.3.0 stays a minor.

What six reviews found

A failed write was reported as success. Node hands end's callback the stream's error; the callback took no arguments and resolved regardless, so on ENOSPC mid-download the command printed done: N rows, recorded complete: true and exited 0 with a truncated file no rerun would continue. The stream also had no 'error' listener until the flush, so an earlier failure became an unhandled 'error' event that killed the run and left five of a destination's six parks unattempted.

A failure on a resumed run deleted every row already downloaded, because written === 0 means "this process wrote nothing", not "the file is empty".

The entry-point guard was wrong for an installed binary. It compared import.meta.url against process.argv[1], and npm installs a bin as a symlink — so npx themeparks-backfill --version printed nothing and exited 0. Green on all 186 unit tests and working when run inside the repo. Only packing and installing the tarball showed it. The runner is its own file now, and scripts/check-package.ts packs, installs and runs the binary in CI.

The vendored OpenAPI schema was stale. unknownMinutes, inParkHours and extremeWaits are on every row the API returns and were in none of the types. The column list is generated from the spec now, and spec-drift.yml commits it alongside the schema — it previously committed only the schema, which would have silently restored the exact defect this release fixes, with nothing red anywhere.

HistoryPage was never exported although the changelog advertised it (dist/index.d.ts had zero references).

Plus: a bare \r in an entity name written unquoted, so one row parsed as two; --list with no value exiting 2 though the help says [TEXT]; no -h; 東京 listing all 127 parks because normalize folds it to ''; an empty --api-key counting as a key; no-arguments fetching /destinations before saying so, which exits 75 with no network; --version printing a bare number; the 404 hint sitting where nothing could reach it; and a unique substring being refused while the Python SDK downloaded it.

The shared contract

test/fixtures/csv_contract.json is asserted by both SDKs. Before it existed this SDK wrote 32 columns and Python wrote 41, with inParkScheduledMinutes against inParkHoursScheduledMinutes, and the test that claimed to check it compared four strings. Both now compute the same fingerprint, 1b6ea478049dc7ca, over the same header.

The state file is <parkId>.<format>.backfill-state.json carrying the SDK, its version, a state version and that fingerprint, refusing anything that does not match — including a state file written by the Python SDK, whose keys differ only on the interrupted path, which is how a cross-SDK resume produced 172 rows where 108 belonged.

Tests

211 (from 186), none needing the network, and 13 of 13 targeted mutations killed — including a renamed column in the generated file, caught by the shared contract. Fixtures are real captures: a real nested 403 body, 101 real destinations, and two real consecutive pages of one request where the newest row, the page's last day and the resume point are three different dates.

One honest gap: history_rate_limited.json is hand-written, and the SDK reads only the Retry-After header, so the fixture's body earns nothing. The README now says so — the hourly-budget header contract is unverified against production, and capturing a real 429 costs an hour of a metered key's budget.

Companion: ThemeParks_Python#29.

🤖 Generated with Claude Code

…d in it

The Python SDK shipped this command first. Porting it, diffing the two outputs
over the same park, and then reviewing the result from six angles found around
twenty defects between them. Two independent implementations reading one API
disagree in exactly the places one of them is wrong.

Magic Kingdom's full five-year archive now comes back byte for byte identical from
both SDKs: 94,223 rows, 41 columns, the only differences being today's row, which
grows as the day elapses.

  npm install themeparks
  npx themeparks-backfill "magic kingdom"

A park or a destination, by name or id, one file per park. NDJSON by default,
--format csv for one wide row per entity per day. Every row carries parkId,
parkName, entityId, entityName and entityType, so two files load into one table
and (entityId, date) is the natural key. The entity name is the one the history
response gave for those rows, not the park's current children list, because rides
get renamed and today's name on a row from three years ago rewrites the record.
Files are named for the park's id, because names change. --list needs no key, so
you can find your park before deciding whether to pay.

Resumable: it checkpoints against the hourly history budget and exits 75, so a
timer retries rather than alerting. The checkpoint is the day the server's own
`next` URL starts on, never the newest row written -- an entity that stopped
reporting has no rows for the tail days of its page, so a row-derived checkpoint
re-fetches days already in the file. That boundary is reported through a new
`onPage` hook on days(), since the page boundary is the server's answer to "where
do I carry on" and the rows cannot tell you.

The state file is <parkId>.<format>.backfill-state.json and records the SDK, its
version, a state version and a fingerprint of the exact header. Anything that does
not match is refused with a message saying why. One state file for two formats
doubled every row on an ndjson -> csv -> ndjson round trip; no record of the header
let a 19-column file resume under a 41-column build; and the two SDKs' keys
differed only on the interrupted path, so the safe paths interoperated and a
cross-SDK resume produced 172 rows where 108 belonged.

Three things only a review caught. A failed write was reported as success: Node
hands `end`'s callback the stream's error and the callback took no arguments, so
ENOSPC mid-download printed "done", recorded complete, and exited 0 with a
truncated file. A failure on a resumed run deleted every row already downloaded,
because `written === 0` means this process wrote nothing rather than the file is
empty. And the executable's entry-point guard compared import.meta.url against
process.argv[1], which npm makes a symlink, so `npx themeparks-backfill --version`
printed nothing and exited 0 -- green on all 186 unit tests and working in the
repo. The runner is its own file now and scripts/check-package.ts packs, installs
and runs the binary in CI.

The vendored OpenAPI schema was stale, so unknownMinutes, inParkHours and
extremeWaits were on every row the API returns and in none of the types. The CSV
column list is generated from the spec, the nightly drift job commits it alongside
the schema, and test/fixtures/csv_contract.json is asserted by both SDKs so the
two headers cannot diverge again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@cubehouse
cubehouse merged commit 486334f into main Sep 28, 2026
3 checks passed
@cubehouse
cubehouse deleted the feat/backfill-parity branch September 28, 2026 17:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant