You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 486334f
Browse filesBrowse the repository at this point in the historyBrowse files
feat(backfill): themeparks-backfill, and the defects six reviews found in it (#51)
The Python SDK shipped this command first. Porting it, diffing the two outputs
over the same park, and then reviewing the result from six angles found around
twenty defects between them. Two independent implementations reading one API
disagree in exactly the places one of them is wrong.
Magic Kingdom's full five-year archive now comes back byte for byte identical from
both SDKs: 94,223 rows, 41 columns, the only differences being today's row, which
grows as the day elapses.
npm install themeparks
npx themeparks-backfill "magic kingdom"
A park or a destination, by name or id, one file per park. NDJSON by default,
--format csv for one wide row per entity per day. Every row carries parkId,
parkName, entityId, entityName and entityType, so two files load into one table
and (entityId, date) is the natural key. The entity name is the one the history
response gave for those rows, not the park's current children list, because rides
get renamed and today's name on a row from three years ago rewrites the record.
Files are named for the park's id, because names change. --list needs no key, so
you can find your park before deciding whether to pay.
Resumable: it checkpoints against the hourly history budget and exits 75, so a
timer retries rather than alerting. The checkpoint is the day the server's own
`next` URL starts on, never the newest row written -- an entity that stopped
reporting has no rows for the tail days of its page, so a row-derived checkpoint
re-fetches days already in the file. That boundary is reported through a new
`onPage` hook on days(), since the page boundary is the server's answer to "where
do I carry on" and the rows cannot tell you.
The state file is <parkId>.<format>.backfill-state.json and records the SDK, its
version, a state version and a fingerprint of the exact header. Anything that does
not match is refused with a message saying why. One state file for two formats
doubled every row on an ndjson -> csv -> ndjson round trip; no record of the header
let a 19-column file resume under a 41-column build; and the two SDKs' keys
differed only on the interrupted path, so the safe paths interoperated and a
cross-SDK resume produced 172 rows where 108 belonged.
Three things only a review caught. A failed write was reported as success: Node
hands `end`'s callback the stream's error and the callback took no arguments, so
ENOSPC mid-download printed "done", recorded complete, and exited 0 with a
truncated file. A failure on a resumed run deleted every row already downloaded,
because `written === 0` means this process wrote nothing rather than the file is
empty. And the executable's entry-point guard compared import.meta.url against
process.argv[1], which npm makes a symlink, so `npx themeparks-backfill --version`
printed nothing and exited 0 -- green on all 186 unit tests and working in the
repo. The runner is its own file now and scripts/check-package.ts packs, installs
and runs the binary in CI.
The vendored OpenAPI schema was stale, so unknownMinutes, inParkHours and
extremeWaits were on every row the API returns and in none of the types. The CSV
column list is generated from the spec, the nightly drift job commits it alongside
the schema, and test/fixtures/csv_contract.json is asserted by both SDKs so the
two headers cannot diverge again.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
0 commit comments