Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,3 +37,6 @@ jobs:
- run: npm run typecheck
- run: npm run build
- run: npm test
# Packs the tarball, installs it elsewhere and runs the binary. The unit
# suite cannot see a bin that does not resolve when installed.
- run: npm run test:package
10 changes: 9 additions & 1 deletion .github/workflows/spec-drift.yml
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,15 @@ jobs:
uses: peter-evans/create-pull-request@v8
with:
token: ${{ steps.app-token.outputs.token }}
add-paths: src/_generated/schema.ts
# BOTH GENERATED FILES. `npm run regenerate` writes the schema and the
# daily column list; committing only the schema meant the next field
# upstream added would be typed and not exported, the CSV would silently
# drop it again, and the Python SDK -- which derives its columns at
# runtime -- would pick it up, so the two headers would diverge with
# nothing red anywhere.
add-paths: |
src/_generated/schema.ts
src/_generated/dailyColumns.ts
branch: automated/spec-drift
commit-message: 'chore: regenerate schema from upstream spec'
title: 'chore: regenerate schema from upstream spec'
Expand Down
91 changes: 91 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,97 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [8.3.0] - 2026-09-28

### Added

- **`themeparks-backfill`: the archive download as a command.** The Python SDK
shipped this first; this is the same tool, and the two write byte-for-byte
identical CSVs. Magic Kingdom's full five-year archive: 94,223 rows, 41 columns,
identical from both, the only differences being today's row, which grows as the
day elapses.

```bash
npm install themeparks
npx themeparks-backfill "magic kingdom"
```

- Takes a park or a **destination**, by name or id, and a name that identifies one
park unambiguously is enough. A destination back fills every park in it, one
file each. An ambiguous name lists the ids that match, sorted by park name.
- `--list [text]` prints destinations with their parks underneath and **needs no
key**, so you can find your park before deciding whether to pay.
- **Runs without a key**, reading the 7 days anonymous access allows, and says
what a key would add.
- NDJSON by default, `--format csv` for one wide row per entity per day. Every row
carries `parkId`, `parkName`, `entityId`, `entityName` and `entityType`, so two
files load into one table and `(entityId, date)` is the natural key. The entity
name is the one the history response gave for those rows, not the park's current
children list: rides get renamed, and today's name on a row from three years ago
rewrites the record. Files are named for the park's id, because names change.
- The CSV carries a **UTF-8 BOM** so Excel on Windows does not mangle `®` and
accents, and a cell a spreadsheet would execute as a formula is prefixed with an
apostrophe. Numeric cells are untouched, so a negative number stays a number.
- **Resumable.** It checkpoints against the hourly history budget and exits 75
(`EX_TEMPFAIL`), so a cron or systemd timer retries rather than alerting, and
running the same command again continues. The checkpoint is the day the server's
own `next` URL starts on, never the newest row written -- an entity that stopped
reporting has no rows for the tail days of its page, so resuming from a row
re-fetches days already in the file.
- The state file is `<parkId>.<format>.backfill-state.json` and records the SDK,
its version, a state version and a fingerprint of the exact header. Anything
that does not match is refused with a message saying why, never resumed --
including a state file written by the Python SDK, whose keys differ.
- One park's failure does not abandon the rest of a destination; what did not
finish is named at the end. A network failure or timeout exits 75, anything the
API rejected exits 1, and neither is a traceback.
- An earlier run's rows are never deleted. A failure or a closed window on a
resumed run keeps the file and says the run did not finish.

- **`onPage` on `days()`**, called once every row of a page has been yielded, with a
`HistoryPage` (`from`, `to`, `next`). The page boundary is the server's own answer
to "where do I carry on", and the rows cannot tell you -- so it is the only safe
checkpoint for a resumable download. `HistoryPage` and `PageOptions` are exported.

- **`DailyEntry` carries `name` and `entityType`**, taken from the history response
itself. Already in the payload, so nothing has to ask what an id refers to. Both
are required fields, so a hand-built `DailyEntry` in a test double needs them.

- **`test/fixtures/csv_contract.json`**, an identical copy of which lives in the
Python SDK. Both suites assert their column list against it, because this is one
command with two implementations and a customer using both should get one file
format. Before it existed, this SDK wrote 32 columns and Python wrote 41.

- **`npm run test:package`**, in CI and `prepublishOnly`: it packs the tarball,
installs it elsewhere and runs the binary. See below for why.

### Fixed

- **The vendored OpenAPI schema was stale.** `unknownMinutes`, `inParkHours` (the
day's numbers limited to the park's published hours) and `extremeWaits` (how many
readings of 480+ minutes are folded into the statistics, which is how you spot a
feed error) are on the rows the API returns and were in none of the types. The CSV
column list is now **generated from the spec**, the nightly drift job commits it
alongside the schema, and the generator refuses a duplicate column name or a
missing nested block.

- **A failed write was reported as success.** Node hands `end`'s callback the
stream's error; the callback took no arguments and resolved regardless, so on
ENOSPC or EDQUOT mid-download the command printed `done: N rows`, recorded
`complete: true` and exited 0 with a truncated file that no rerun would continue.
The stream also had no `'error'` listener until the flush, so an earlier failure
became an unhandled `'error'` event that killed the whole run.

- **A bare `\r` in an entity name was written unquoted**, so one row parsed as two
with every later column shifted.

- **`--list` with no value exited 2** although the help advertises `--list [TEXT]`;
`-h` was not accepted; a query that folds to nothing (`東京`) listed all 127 parks
instead of none; an empty `--api-key` or `THEMEPARKS_API_KEY=""` counted as a key;
running with no arguments fetched `/destinations` before saying so, which exited
75 with no network; `--version` printed a bare number; and the 404 hint for a
mistyped id sat where nothing could reach it.

## [8.2.0] - 2026-09-26

### Added
Expand Down
28 changes: 27 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -332,7 +332,33 @@ try {
`history.changeRows(query)` is the same treatment for `changes`: one flattened
stream of `{ entityId, row }` whether you asked a park or a ride.

A complete backfill with resume and CSV output is in
### Or skip the code: there is a command

Installing the package puts `themeparks-backfill` on your path. It is the same
job as the example below, resumable, and it is what to reach for if what you
want is the file rather than the code:

```bash
npx themeparks-backfill --list disney # find your park. No key needed.
npx themeparks-backfill "magic kingdom" # NDJSON, into the current directory
npx themeparks-backfill "Walt Disney World Resort" --format csv --out ./data
```

A park or a **destination**, by name or by id; a destination writes one file per
park. Every row carries `parkId`, `parkName`, `entityId`, `entityName` and
`entityType`, so two files load into one table and `(entityId, date)` is the
natural key. Files are named for the park's id, because names change.

How far back it reaches is your plan, and it asks the API rather than making you
work it out. It checkpoints against the hourly history budget and exits 75
(`EX_TEMPFAIL`) when that runs out, so a cron or systemd timer retries instead of
alerting and the same command continues where it stopped. `--help` has the rest.

The CSV is byte-for-byte identical to the Python SDK's, which runs the same
command: Magic Kingdom's five-year archive is 94,223 rows and 41 columns from
either.

A library version of the same loop, if you want to own it, is in
[`examples/backfill.mjs`](examples/backfill.mjs). It pulled Disneyland Resort's
whole daily archive, 98,452 rows, in one run.

Expand Down
10 changes: 7 additions & 3 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

11 changes: 8 additions & 3 deletions package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "themeparks",
"version": "8.2.0",
"version": "8.3.0",
"description": "Official SDK for the ThemeParks.wiki API",
"license": "MIT",
"repository": "github:ThemeParks/ThemeParks_JavaScript",
Expand All @@ -17,6 +17,9 @@
"require": "./dist/index.cjs"
}
},
"bin": {
"themeparks-backfill": "./dist/backfill-cli.js"
},
"files": [
"dist",
"README.md",
Expand All @@ -39,7 +42,8 @@
"regenerate": "tsx scripts/regenerate.ts",
"docs": "typedoc",
"docs:serve": "npx serve docs-site",
"prepublishOnly": "npm run build"
"prepublishOnly": "npm run build && npm run test:package",
"test:package": "tsx scripts/check-package.ts"
},
"devDependencies": {
"@types/node": "^26.4.1",
Expand All @@ -53,6 +57,7 @@
"typedoc": "^0.28.19",
"typedoc-plugin-markdown": "^4.11.0",
"typescript": "^5.4.0",
"vitest": "^4.1.11"
"vitest": "^4.1.11",
"yaml": "^2.5.0"
}
}
66 changes: 66 additions & 0 deletions scripts/check-package.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
/**
* Pack the tarball, install it somewhere else, and run the binary.
*
* This exists because of a defect that passed the whole unit suite, worked in the
* repo, and would have shipped: the executable's "am I the entry point" guard
* compared `import.meta.url` against `process.argv[1]`, and npm installs a binary
* as a SYMLINK in `node_modules/.bin`. The paths differ, the guard was false, and
* `npx themeparks-backfill --version` printed nothing and exited 0.
*
* Nothing short of installing it shows that. So: pack, install into a temporary
* directory, run the binary the way a customer does, and require real output.
*/
import { execFileSync } from 'node:child_process';
import { mkdtempSync, readdirSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';

const root = process.cwd();
const dir = mkdtempSync(join(tmpdir(), 'themeparks-package-'));

function sh(command: string, args: string[], cwd: string): string {
return execFileSync(command, args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] });
}

try {
sh('npm', ['pack', '--pack-destination', dir], root);
const tarball = readdirSync(dir).find((f) => f.endsWith('.tgz'));
if (tarball === undefined) throw new Error('npm pack produced no tarball');

writeFileSync(join(dir, 'package.json'), JSON.stringify({ name: 'consumer', private: true }));
sh('npm', ['install', '--no-audit', '--no-fund', join(dir, tarball)], dir);

const bin = join(dir, 'node_modules', '.bin', 'themeparks-backfill');
const version = sh(bin, ['--version'], dir).trim();
const expected = JSON.parse(sh('npm', ['pkg', 'get', 'version'], root) as string) as string;
// Named, matching the Python SDK: a bare number cannot be pasted into a bug report.
if (version !== `themeparks-backfill ${expected}`) {
throw new Error(
`the installed binary printed "${version}", expected "themeparks-backfill ${expected}"`,
);
}

const help = sh(bin, ['--help'], dir);
if (!help.includes('themeparks-backfill')) throw new Error('--help printed nothing usable');

// A bare `--list` exits 2 if parseArgs treats the option as value-taking, which is
// exactly the class of defect only an installed run shows. Needs no key.
const listed = sh(bin, ['--list', 'epcot'], dir);
if (!listed.includes('EPCOT')) throw new Error(`--list epcot printed nothing usable: ${listed}`);

// The library import path, which is a different resolution from the binary.
const imported = sh(
process.execPath,
[
'--input-type=module',
'-e',
"import {ThemeParks} from 'themeparks'; console.log(typeof ThemeParks)",
],
dir,
).trim();
if (imported !== 'function') throw new Error(`importing the package gave ${imported}`);

console.log(`package ok: binary and import both work from a clean install (${version})`);
} finally {
rmSync(dir, { recursive: true, force: true });
}
71 changes: 71 additions & 0 deletions scripts/regenerate.ts
Original file line number Diff line number Diff line change
@@ -1,9 +1,11 @@
import { writeFile } from 'node:fs/promises';
import { resolve } from 'node:path';
import openapiTS, { astToString } from 'openapi-typescript';
import { parse } from 'yaml';

const SPEC_URL = 'https://api.themeparks.wiki/docs/v1.yaml';
const OUTPUT = resolve(process.cwd(), 'src/_generated/schema.ts');
const COLUMNS_OUTPUT = resolve(process.cwd(), 'src/_generated/dailyColumns.ts');

const header = `/* eslint-disable */
/**
Expand All @@ -12,11 +14,80 @@ const header = `/* eslint-disable */
*/
`;

interface SchemaNode {
properties?: Record<string, { $ref?: string; type?: string }>;
$ref?: string;
}

/**
* Every scalar on a daily history row, flattened to one column name each, in the
* order the spec declares them.
*
* GENERATED, not typed out. The hand-written list in backfill.ts had drifted
* three ways at once: `unknownMinutes` and the whole `inParkHours` block were on
* every row the API returns and in no column, `extremeWaits` likewise, and
* `singleRider` carried two of its five percentiles while `standby` carried all
* five. Ten of thirty-six fields silently absent from a file people pay for, and
* the two SDKs disagreeing about the header of a file they both claim to write.
* Regenerate and the columns follow.
*/
function dailyColumns(schemas: Record<string, SchemaNode>): string[] {
const walk = (name: string, prefix: string): string[] => {
const node = schemas[name];
const out: string[] = [];
for (const [field, value] of Object.entries(node?.properties ?? {})) {
const head = prefix === '' ? field : `${prefix}${field[0]!.toUpperCase()}${field.slice(1)}`;
const ref = value.$ref?.split('/').pop();
if (ref !== undefined && schemas[ref]?.properties !== undefined) {
out.push(...walk(ref, head));
} else {
out.push(head);
}
}
return out;
};
return walk('HistoryDailyRow', '');
}

async function main() {
const ast = await openapiTS(new URL(SPEC_URL));
const body = astToString(ast);
await writeFile(OUTPUT, header + body, 'utf8');
console.log(`Wrote ${OUTPUT}`);

const spec = parse(await (await fetch(SPEC_URL)).text()) as {
components: { schemas: Record<string, SchemaNode> };
};
const columns = dailyColumns(spec.components.schemas);
// A FLOOR NEAR THE REAL COUNT. `< 10` caught a total wipe-out and nothing else:
// a spec that expressed one nested block inline or behind allOf would drop seven
// columns and pass, and the header would gain an always-empty `inParkHours`.
if (columns.length < 30) {
throw new Error(
`only ${String(columns.length)} daily columns, expected 36ish: has the spec moved to allOf/inline blocks?`,
);
}
for (const block of ['standby', 'singleRider', 'extremeWaits', 'inParkHours']) {
if (!columns.some((c) => c.startsWith(block))) {
throw new Error(
`no ${block} columns: the generator only follows $ref, and this block is no longer one`,
);
}
}
// The <outer><Inner> rule is not injective. A collision writes one value into two
// slots under a right-looking header, so it fails the build instead.
const dupes = columns.filter((c, i) => columns.indexOf(c) !== i);
if (dupes.length > 0) throw new Error(`duplicate daily column names: ${dupes.join(', ')}`);
// No eslint-disable on this one: it is a plain array, and an unused directive
// is itself a warning.
const columnsHeader = header.replace('/* eslint-disable */\n', '');
await writeFile(
COLUMNS_OUTPUT,
`${columnsHeader}\n/** Every scalar on a daily history row, flattened, in spec order. */\n` +
`export const DAILY_COLUMNS = [\n${columns.map((c) => ` '${c}',`).join('\n')}\n] as const;\n`,
'utf8',
);
console.log(`Wrote ${COLUMNS_OUTPUT} (${columns.length} columns)`);
}

main().catch((err) => {
Expand Down
Loading
Loading