Skip to content

Commit 486334f

Browse files
cubehouseclaude
andauthored
feat(backfill): themeparks-backfill, and the defects six reviews found in it (#51)
The Python SDK shipped this command first. Porting it, diffing the two outputs over the same park, and then reviewing the result from six angles found around twenty defects between them. Two independent implementations reading one API disagree in exactly the places one of them is wrong. Magic Kingdom's full five-year archive now comes back byte for byte identical from both SDKs: 94,223 rows, 41 columns, the only differences being today's row, which grows as the day elapses. npm install themeparks npx themeparks-backfill "magic kingdom" A park or a destination, by name or id, one file per park. NDJSON by default, --format csv for one wide row per entity per day. Every row carries parkId, parkName, entityId, entityName and entityType, so two files load into one table and (entityId, date) is the natural key. The entity name is the one the history response gave for those rows, not the park's current children list, because rides get renamed and today's name on a row from three years ago rewrites the record. Files are named for the park's id, because names change. --list needs no key, so you can find your park before deciding whether to pay. Resumable: it checkpoints against the hourly history budget and exits 75, so a timer retries rather than alerting. The checkpoint is the day the server's own `next` URL starts on, never the newest row written -- an entity that stopped reporting has no rows for the tail days of its page, so a row-derived checkpoint re-fetches days already in the file. That boundary is reported through a new `onPage` hook on days(), since the page boundary is the server's answer to "where do I carry on" and the rows cannot tell you. The state file is <parkId>.<format>.backfill-state.json and records the SDK, its version, a state version and a fingerprint of the exact header. Anything that does not match is refused with a message saying why. One state file for two formats doubled every row on an ndjson -> csv -> ndjson round trip; no record of the header let a 19-column file resume under a 41-column build; and the two SDKs' keys differed only on the interrupted path, so the safe paths interoperated and a cross-SDK resume produced 172 rows where 108 belonged. Three things only a review caught. A failed write was reported as success: Node hands `end`'s callback the stream's error and the callback took no arguments, so ENOSPC mid-download printed "done", recorded complete, and exited 0 with a truncated file. A failure on a resumed run deleted every row already downloaded, because `written === 0` means this process wrote nothing rather than the file is empty. And the executable's entry-point guard compared import.meta.url against process.argv[1], which npm makes a symlink, so `npx themeparks-backfill --version` printed nothing and exited 0 -- green on all 186 unit tests and working in the repo. The runner is its own file now and scripts/check-package.ts packs, installs and runs the binary in CI. The vendored OpenAPI schema was stale, so unknownMinutes, inParkHours and extremeWaits were on every row the API returns and in none of the types. The CSV column list is generated from the spec, the nightly drift job commits it alongside the schema, and test/fixtures/csv_contract.json is asserted by both SDKs so the two headers cannot diverge again. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent f53397c commit 486334f

23 files changed

Lines changed: 5780 additions & 184 deletions

‎.github/workflows/ci.yml‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -37,3 +37,6 @@ jobs:
3737
- run: npm run typecheck
3838
- run: npm run build
3939
- run: npm test
40+
# Packs the tarball, installs it elsewhere and runs the binary. The unit
41+
# suite cannot see a bin that does not resolve when installed.
42+
- run: npm run test:package

‎.github/workflows/spec-drift.yml‎

Lines changed: 9 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -58,7 +58,15 @@ jobs:
5858
uses: peter-evans/create-pull-request@v8
5959
with:
6060
token: ${{ steps.app-token.outputs.token }}
61-
add-paths: src/_generated/schema.ts
61+
# BOTH GENERATED FILES. `npm run regenerate` writes the schema and the
62+
# daily column list; committing only the schema meant the next field
63+
# upstream added would be typed and not exported, the CSV would silently
64+
# drop it again, and the Python SDK -- which derives its columns at
65+
# runtime -- would pick it up, so the two headers would diverge with
66+
# nothing red anywhere.
67+
add-paths: |
68+
src/_generated/schema.ts
69+
src/_generated/dailyColumns.ts
6270
branch: automated/spec-drift
6371
commit-message: 'chore: regenerate schema from upstream spec'
6472
title: 'chore: regenerate schema from upstream spec'

‎CHANGELOG.md‎

Lines changed: 91 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -5,6 +5,97 @@ All notable changes to this project will be documented in this file.
55
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
66
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
77

8+
## [8.3.0] - 2026-09-28
9+
10+
### Added
11+
12+
- **`themeparks-backfill`: the archive download as a command.** The Python SDK
13+
shipped this first; this is the same tool, and the two write byte-for-byte
14+
identical CSVs. Magic Kingdom's full five-year archive: 94,223 rows, 41 columns,
15+
identical from both, the only differences being today's row, which grows as the
16+
day elapses.
17+
18+
```bash
19+
npm install themeparks
20+
npx themeparks-backfill "magic kingdom"
21+
```
22+
23+
- Takes a park or a **destination**, by name or id, and a name that identifies one
24+
park unambiguously is enough. A destination back fills every park in it, one
25+
file each. An ambiguous name lists the ids that match, sorted by park name.
26+
- `--list [text]` prints destinations with their parks underneath and **needs no
27+
key**, so you can find your park before deciding whether to pay.
28+
- **Runs without a key**, reading the 7 days anonymous access allows, and says
29+
what a key would add.
30+
- NDJSON by default, `--format csv` for one wide row per entity per day. Every row
31+
carries `parkId`, `parkName`, `entityId`, `entityName` and `entityType`, so two
32+
files load into one table and `(entityId, date)` is the natural key. The entity
33+
name is the one the history response gave for those rows, not the park's current
34+
children list: rides get renamed, and today's name on a row from three years ago
35+
rewrites the record. Files are named for the park's id, because names change.
36+
- The CSV carries a **UTF-8 BOM** so Excel on Windows does not mangle `®` and
37+
accents, and a cell a spreadsheet would execute as a formula is prefixed with an
38+
apostrophe. Numeric cells are untouched, so a negative number stays a number.
39+
- **Resumable.** It checkpoints against the hourly history budget and exits 75
40+
(`EX_TEMPFAIL`), so a cron or systemd timer retries rather than alerting, and
41+
running the same command again continues. The checkpoint is the day the server's
42+
own `next` URL starts on, never the newest row written -- an entity that stopped
43+
reporting has no rows for the tail days of its page, so resuming from a row
44+
re-fetches days already in the file.
45+
- The state file is `<parkId>.<format>.backfill-state.json` and records the SDK,
46+
its version, a state version and a fingerprint of the exact header. Anything
47+
that does not match is refused with a message saying why, never resumed --
48+
including a state file written by the Python SDK, whose keys differ.
49+
- One park's failure does not abandon the rest of a destination; what did not
50+
finish is named at the end. A network failure or timeout exits 75, anything the
51+
API rejected exits 1, and neither is a traceback.
52+
- An earlier run's rows are never deleted. A failure or a closed window on a
53+
resumed run keeps the file and says the run did not finish.
54+
55+
- **`onPage` on `days()`**, called once every row of a page has been yielded, with a
56+
`HistoryPage` (`from`, `to`, `next`). The page boundary is the server's own answer
57+
to "where do I carry on", and the rows cannot tell you -- so it is the only safe
58+
checkpoint for a resumable download. `HistoryPage` and `PageOptions` are exported.
59+
60+
- **`DailyEntry` carries `name` and `entityType`**, taken from the history response
61+
itself. Already in the payload, so nothing has to ask what an id refers to. Both
62+
are required fields, so a hand-built `DailyEntry` in a test double needs them.
63+
64+
- **`test/fixtures/csv_contract.json`**, an identical copy of which lives in the
65+
Python SDK. Both suites assert their column list against it, because this is one
66+
command with two implementations and a customer using both should get one file
67+
format. Before it existed, this SDK wrote 32 columns and Python wrote 41.
68+
69+
- **`npm run test:package`**, in CI and `prepublishOnly`: it packs the tarball,
70+
installs it elsewhere and runs the binary. See below for why.
71+
72+
### Fixed
73+
74+
- **The vendored OpenAPI schema was stale.** `unknownMinutes`, `inParkHours` (the
75+
day's numbers limited to the park's published hours) and `extremeWaits` (how many
76+
readings of 480+ minutes are folded into the statistics, which is how you spot a
77+
feed error) are on the rows the API returns and were in none of the types. The CSV
78+
column list is now **generated from the spec**, the nightly drift job commits it
79+
alongside the schema, and the generator refuses a duplicate column name or a
80+
missing nested block.
81+
82+
- **A failed write was reported as success.** Node hands `end`'s callback the
83+
stream's error; the callback took no arguments and resolved regardless, so on
84+
ENOSPC or EDQUOT mid-download the command printed `done: N rows`, recorded
85+
`complete: true` and exited 0 with a truncated file that no rerun would continue.
86+
The stream also had no `'error'` listener until the flush, so an earlier failure
87+
became an unhandled `'error'` event that killed the whole run.
88+
89+
- **A bare `\r` in an entity name was written unquoted**, so one row parsed as two
90+
with every later column shifted.
91+
92+
- **`--list` with no value exited 2** although the help advertises `--list [TEXT]`;
93+
`-h` was not accepted; a query that folds to nothing (`東京`) listed all 127 parks
94+
instead of none; an empty `--api-key` or `THEMEPARKS_API_KEY=""` counted as a key;
95+
running with no arguments fetched `/destinations` before saying so, which exited
96+
75 with no network; `--version` printed a bare number; and the 404 hint for a
97+
mistyped id sat where nothing could reach it.
98+
899
## [8.2.0] - 2026-09-26
9100

10101
### Added

‎README.md‎

Lines changed: 27 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -332,7 +332,33 @@ try {
332332
`history.changeRows(query)` is the same treatment for `changes`: one flattened
333333
stream of `{ entityId, row }` whether you asked a park or a ride.
334334
335-
A complete backfill with resume and CSV output is in
335+
### Or skip the code: there is a command
336+
337+
Installing the package puts `themeparks-backfill` on your path. It is the same
338+
job as the example below, resumable, and it is what to reach for if what you
339+
want is the file rather than the code:
340+
341+
```bash
342+
npx themeparks-backfill --list disney # find your park. No key needed.
343+
npx themeparks-backfill "magic kingdom" # NDJSON, into the current directory
344+
npx themeparks-backfill "Walt Disney World Resort" --format csv --out ./data
345+
```
346+
347+
A park or a **destination**, by name or by id; a destination writes one file per
348+
park. Every row carries `parkId`, `parkName`, `entityId`, `entityName` and
349+
`entityType`, so two files load into one table and `(entityId, date)` is the
350+
natural key. Files are named for the park's id, because names change.
351+
352+
How far back it reaches is your plan, and it asks the API rather than making you
353+
work it out. It checkpoints against the hourly history budget and exits 75
354+
(`EX_TEMPFAIL`) when that runs out, so a cron or systemd timer retries instead of
355+
alerting and the same command continues where it stopped. `--help` has the rest.
356+
357+
The CSV is byte-for-byte identical to the Python SDK's, which runs the same
358+
command: Magic Kingdom's five-year archive is 94,223 rows and 41 columns from
359+
either.
360+
361+
A library version of the same loop, if you want to own it, is in
336362
[`examples/backfill.mjs`](examples/backfill.mjs). It pulled Disneyland Resort's
337363
whole daily archive, 98,452 rows, in one run.
338364

‎package-lock.json‎

Lines changed: 7 additions & 3 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.

‎package.json‎

Lines changed: 8 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
{
22
"name": "themeparks",
3-
"version": "8.2.0",
3+
"version": "8.3.0",
44
"description": "Official SDK for the ThemeParks.wiki API",
55
"license": "MIT",
66
"repository": "github:ThemeParks/ThemeParks_JavaScript",
@@ -17,6 +17,9 @@
1717
"require": "./dist/index.cjs"
1818
}
1919
},
20+
"bin": {
21+
"themeparks-backfill": "./dist/backfill-cli.js"
22+
},
2023
"files": [
2124
"dist",
2225
"README.md",
@@ -39,7 +42,8 @@
3942
"regenerate": "tsx scripts/regenerate.ts",
4043
"docs": "typedoc",
4144
"docs:serve": "npx serve docs-site",
42-
"prepublishOnly": "npm run build"
45+
"prepublishOnly": "npm run build && npm run test:package",
46+
"test:package": "tsx scripts/check-package.ts"
4347
},
4448
"devDependencies": {
4549
"@types/node": "^26.4.1",
@@ -53,6 +57,7 @@
5357
"typedoc": "^0.28.19",
5458
"typedoc-plugin-markdown": "^4.11.0",
5559
"typescript": "^5.4.0",
56-
"vitest": "^4.1.11"
60+
"vitest": "^4.1.11",
61+
"yaml": "^2.5.0"
5762
}
5863
}

‎scripts/check-package.ts‎

Lines changed: 66 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,66 @@
1+
/**
2+
* Pack the tarball, install it somewhere else, and run the binary.
3+
*
4+
* This exists because of a defect that passed the whole unit suite, worked in the
5+
* repo, and would have shipped: the executable's "am I the entry point" guard
6+
* compared `import.meta.url` against `process.argv[1]`, and npm installs a binary
7+
* as a SYMLINK in `node_modules/.bin`. The paths differ, the guard was false, and
8+
* `npx themeparks-backfill --version` printed nothing and exited 0.
9+
*
10+
* Nothing short of installing it shows that. So: pack, install into a temporary
11+
* directory, run the binary the way a customer does, and require real output.
12+
*/
13+
import { execFileSync } from 'node:child_process';
14+
import { mkdtempSync, readdirSync, rmSync, writeFileSync } from 'node:fs';
15+
import { tmpdir } from 'node:os';
16+
import { join } from 'node:path';
17+
18+
const root = process.cwd();
19+
const dir = mkdtempSync(join(tmpdir(), 'themeparks-package-'));
20+
21+
function sh(command: string, args: string[], cwd: string): string {
22+
return execFileSync(command, args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] });
23+
}
24+
25+
try {
26+
sh('npm', ['pack', '--pack-destination', dir], root);
27+
const tarball = readdirSync(dir).find((f) => f.endsWith('.tgz'));
28+
if (tarball === undefined) throw new Error('npm pack produced no tarball');
29+
30+
writeFileSync(join(dir, 'package.json'), JSON.stringify({ name: 'consumer', private: true }));
31+
sh('npm', ['install', '--no-audit', '--no-fund', join(dir, tarball)], dir);
32+
33+
const bin = join(dir, 'node_modules', '.bin', 'themeparks-backfill');
34+
const version = sh(bin, ['--version'], dir).trim();
35+
const expected = JSON.parse(sh('npm', ['pkg', 'get', 'version'], root) as string) as string;
36+
// Named, matching the Python SDK: a bare number cannot be pasted into a bug report.
37+
if (version !== `themeparks-backfill ${expected}`) {
38+
throw new Error(
39+
`the installed binary printed "${version}", expected "themeparks-backfill ${expected}"`,
40+
);
41+
}
42+
43+
const help = sh(bin, ['--help'], dir);
44+
if (!help.includes('themeparks-backfill')) throw new Error('--help printed nothing usable');
45+
46+
// A bare `--list` exits 2 if parseArgs treats the option as value-taking, which is
47+
// exactly the class of defect only an installed run shows. Needs no key.
48+
const listed = sh(bin, ['--list', 'epcot'], dir);
49+
if (!listed.includes('EPCOT')) throw new Error(`--list epcot printed nothing usable: ${listed}`);
50+
51+
// The library import path, which is a different resolution from the binary.
52+
const imported = sh(
53+
process.execPath,
54+
[
55+
'--input-type=module',
56+
'-e',
57+
"import {ThemeParks} from 'themeparks'; console.log(typeof ThemeParks)",
58+
],
59+
dir,
60+
).trim();
61+
if (imported !== 'function') throw new Error(`importing the package gave ${imported}`);
62+
63+
console.log(`package ok: binary and import both work from a clean install (${version})`);
64+
} finally {
65+
rmSync(dir, { recursive: true, force: true });
66+
}

‎scripts/regenerate.ts‎

Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,11 @@
11
import { writeFile } from 'node:fs/promises';
22
import { resolve } from 'node:path';
33
import openapiTS, { astToString } from 'openapi-typescript';
4+
import { parse } from 'yaml';
45

56
const SPEC_URL = 'https://api.themeparks.wiki/docs/v1.yaml';
67
const OUTPUT = resolve(process.cwd(), 'src/_generated/schema.ts');
8+
const COLUMNS_OUTPUT = resolve(process.cwd(), 'src/_generated/dailyColumns.ts');
79

810
const header = `/* eslint-disable */
911
/**
@@ -12,11 +14,80 @@ const header = `/* eslint-disable */
1214
*/
1315
`;
1416

17+
interface SchemaNode {
18+
properties?: Record<string, { $ref?: string; type?: string }>;
19+
$ref?: string;
20+
}
21+
22+
/**
23+
* Every scalar on a daily history row, flattened to one column name each, in the
24+
* order the spec declares them.
25+
*
26+
* GENERATED, not typed out. The hand-written list in backfill.ts had drifted
27+
* three ways at once: `unknownMinutes` and the whole `inParkHours` block were on
28+
* every row the API returns and in no column, `extremeWaits` likewise, and
29+
* `singleRider` carried two of its five percentiles while `standby` carried all
30+
* five. Ten of thirty-six fields silently absent from a file people pay for, and
31+
* the two SDKs disagreeing about the header of a file they both claim to write.
32+
* Regenerate and the columns follow.
33+
*/
34+
function dailyColumns(schemas: Record<string, SchemaNode>): string[] {
35+
const walk = (name: string, prefix: string): string[] => {
36+
const node = schemas[name];
37+
const out: string[] = [];
38+
for (const [field, value] of Object.entries(node?.properties ?? {})) {
39+
const head = prefix === '' ? field : `${prefix}${field[0]!.toUpperCase()}${field.slice(1)}`;
40+
const ref = value.$ref?.split('/').pop();
41+
if (ref !== undefined && schemas[ref]?.properties !== undefined) {
42+
out.push(...walk(ref, head));
43+
} else {
44+
out.push(head);
45+
}
46+
}
47+
return out;
48+
};
49+
return walk('HistoryDailyRow', '');
50+
}
51+
1552
async function main() {
1653
const ast = await openapiTS(new URL(SPEC_URL));
1754
const body = astToString(ast);
1855
await writeFile(OUTPUT, header + body, 'utf8');
1956
console.log(`Wrote ${OUTPUT}`);
57+
58+
const spec = parse(await (await fetch(SPEC_URL)).text()) as {
59+
components: { schemas: Record<string, SchemaNode> };
60+
};
61+
const columns = dailyColumns(spec.components.schemas);
62+
// A FLOOR NEAR THE REAL COUNT. `< 10` caught a total wipe-out and nothing else:
63+
// a spec that expressed one nested block inline or behind allOf would drop seven
64+
// columns and pass, and the header would gain an always-empty `inParkHours`.
65+
if (columns.length < 30) {
66+
throw new Error(
67+
`only ${String(columns.length)} daily columns, expected 36ish: has the spec moved to allOf/inline blocks?`,
68+
);
69+
}
70+
for (const block of ['standby', 'singleRider', 'extremeWaits', 'inParkHours']) {
71+
if (!columns.some((c) => c.startsWith(block))) {
72+
throw new Error(
73+
`no ${block} columns: the generator only follows $ref, and this block is no longer one`,
74+
);
75+
}
76+
}
77+
// The <outer><Inner> rule is not injective. A collision writes one value into two
78+
// slots under a right-looking header, so it fails the build instead.
79+
const dupes = columns.filter((c, i) => columns.indexOf(c) !== i);
80+
if (dupes.length > 0) throw new Error(`duplicate daily column names: ${dupes.join(', ')}`);
81+
// No eslint-disable on this one: it is a plain array, and an unused directive
82+
// is itself a warning.
83+
const columnsHeader = header.replace('/* eslint-disable */\n', '');
84+
await writeFile(
85+
COLUMNS_OUTPUT,
86+
`${columnsHeader}\n/** Every scalar on a daily history row, flattened, in spec order. */\n` +
87+
`export const DAILY_COLUMNS = [\n${columns.map((c) => ` '${c}',`).join('\n')}\n] as const;\n`,
88+
'utf8',
89+
);
90+
console.log(`Wrote ${COLUMNS_OUTPUT} (${columns.length} columns)`);
2091
}
2192

2293
main().catch((err) => {

0 commit comments

Comments
 (0)