Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .prettierignore
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,5 @@ docs-site/
coverage/
src/_generated/
*.log
# Shared byte for byte with the Python SDK, which tests against the same capture.
test/fixtures/space_mountain_*.json
116 changes: 116 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,122 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

To be released as **8.4.0**, a minor: the additions below are backwards
compatible at runtime. One is not quite so at the type level: `HistorySpan`
gains a required field, `finalThrough`, so code that builds a `HistorySpan`
object by hand (a test double, say) needs to add it. Code that only reads the
spans `span()` returns is unaffected.

### Added

- **`themeparks-backfill --since YYYY-MM-DD` and `--until YYYY-MM-DD`.** There
was no way to ask for less than everything: `--since` was an unknown option,
so a key that reaches the whole archive downloaded all of it, every time. Both
days are inclusive and must be real calendar days written `YYYY-MM-DD`
(`2025-02-30`, `2025-1-1` and `20250101` are refused), and `--since` after
`--until` is refused before anything is requested. `--since` applies when a
file is started. A later run accepts the same `--since` or a later one, so a
fixed or a rolling cron line both work, and refuses one earlier than the day
the file was started from, or one that would leave a gap, with what to do
instead of quietly handing back a file that is not what was asked for.

- **`history.changeRows()` exposes the `opening` state.** The raw history
response carries, per entity, the state in force at the start of the range,
and `changeRows()` threw it away. Without it the time between midnight and an
entity's first change had no known status, so a day rebuilt from raw history
disagreed with the daily summary whenever a ride was still running from the
night before. The result is still the same async generator and yields exactly
what it always did; it now also has `opening`, an object of `HistoryOpening`
keyed by entity id, covering every entity in the response, including one that
did not change all day. It is readable once the response has arrived: iterate
first, or `await changes.load()`. Either way it is one request. It is not
enumerable, so spreading or logging the result is safe, and after a failed
request it says so. The `HistoryChanges` and `HistoryOpening` types are
exported, and the README explains `opening.degraded`.

- **`HistorySpan.finalThrough`**: the newest day the archive has recorded that
the key may read, the earlier of `recordedTo` and `retrievableThrough`: the
place to stop if you fetch each day once.

- **An interrupted `themeparks-backfill` never appends a day twice.** The state
file is written before the first request and after every page, atomically,
with the size of the data file at that moment; the next run first cuts the
file back to that size. Ctrl-C, SIGTERM (now exit 143, as well as Ctrl-C's 130) and a kill at any point cost at most the page in flight. Two runs on the
same park and `--out` at once are refused.

- **`days()` awaits `onPage`.** A hook that returns a promise holds the next page
until it settles, so a checkpoint written there is on disk before the next
request. The hook's type is widened to return `unknown`, so existing callbacks
still compile.

### Fixed

- **A finished park now updates on the next run.** A rerun printed
`already complete` and exited 0 without fetching a single new day, so a nightly
cron looked healthy and never updated; the only way to get yesterday was
`--overwrite`, which downloads the whole archive again. A finished file is now
carried forward from the day after its last one, appending only the new days.
An interrupted run still resumes at its page boundary, including a rerun
interrupted before its first page, which would otherwise have started again
from the top of the file's range.

- **The newest rows of a backfill were partial days, and stayed that way.** A run
ended on `retrievableThrough`, which is usually today: today's row is the day
so far, and the archive records days 2 to 3 behind live data, so the last few
days of every file were still changing when they were written. A run now ends
at `finalThrough`, says so when it holds days back, and the next run adds them
once they are final. Each day is fetched once, as the archive recorded it; the
README says how to fetch a range again if the archive later re-records it.

**Files written by 8.3.x are corrected once.** Their state file does not say
which of their newest days were final, so the first run of this version
removes the rows dated within seven days of that run's end and fetches those
days again. Every other row is left byte for byte as it was, and a file whose
newest row is older than that is not rewritten at all. A file that lies wholly
inside those seven days, as an anonymous run's does, is downloaded again. The
state file format moves to version 2 for this; version 1 files from this SDK
are upgraded, not refused.

- **A continued file no longer jumps forward to the key's first day.** When the
day a file continues from is older than the key may read (a cron that did not
run for longer than the key's window, or a key that lost its plan), the run
used to carry on from the key's first day and leave a gap the file then did not
record. It is refused now, the file is left alone, and the message says how to
start again.

- **A finished file written to a different column layout was appended to.** Only
an unfinished one was refused. A finished one fell through to a fresh start,
which opened the existing file in append mode and wrote the whole archive into
it a second time under a second header, exit 0. It is refused now, the same
way.

- **An interrupted nightly extension appended the same days twice.** The state
was written only at the end of a run or on an error the command caught, so
Ctrl-C, SIGTERM or a kill during an extension left it saying finished through
the old day, and the rerun appended those days again. See the checkpointing
above.

- **A run that stopped part-way through its first page lost days.** It resumed
from the newest day any entity had reached; rows arrive entity by entity, so
the entities behind it lost the days in between. Checkpoints make this
impossible for new files, and an older state with only that day goes back a
whole page (31 days) instead.

- **The state recorded the start asked for, not the first day written.** `start`
is now the first day the file covers, after the key's window, and a new
`since` field keeps the start that was asked for, so the same `--since` keeps
working and a different one before the file is refused.

- **The anonymous-access notice promised 7 days.** Only final days are written,
so an anonymous run writes usually 4 or 5 days per park, and the notice now
says so.

- **A state file whose data file had been deleted was continued**, producing a
file that started part-way through its range and was then recorded as
complete. The park is downloaded again from the start instead.

## [8.3.1] - 2026-09-28

### Fixed
Expand Down
76 changes: 71 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -265,8 +265,9 @@ Three things to know before polling these:
fails fast by default rather than looking hung. A REST 429, which asks for
seconds, is still ridden out.
- **Today is not final.** The default cache leaves `changes` and `daily`
uncached and keeps `coverage` for an hour. A completed day never changes, so
cache it yourself for as long as you like.
uncached and keeps `coverage` for an hour. A day on or before `recordedTo`
has been recorded, so cache it yourself for as long as you like; the archive
only re-records a past day in the rare case of a feed repair.

`tp.raw.getEntityHistory(id, query)`, `getEntityHistoryDaily(id, query)` and
`getEntityHistoryCoverage(id)` are the underlying calls.
Expand All @@ -283,10 +284,11 @@ const DISNEYLAND = '7340550b-c14d-4def-80bb-acdb51d49a66';
const tp = new ThemeParks({ apiKey: process.env.THEMEPARKS_API_KEY });
const history = tp.entity(DISNEYLAND).history;

// What exists, and what your key may read. The same three fields whether the
// id is a park or a single ride.
// What exists, and what your key may read. The same fields whether the id is a
// park or a single ride.
const span = await history.span();
// -> { archiveFrom: '2021-07-03', recordedTo: '2026-09-22', retrievableThrough: '2026-09-23' }
// -> { archiveFrom: '2021-07-03', recordedTo: '2026-09-22',
// retrievableThrough: '2026-09-23', finalThrough: '2026-09-22' }

// Pages until the server stops offering a `next`, yielding as it goes.
for await (const { entityId, row } of history.days({
Expand All @@ -308,6 +310,14 @@ what your key may read, the second is what the archive holds. They differ on
every plan below the top one, and asking past the entitlement is how a long run
ends in 403s.

**If you write each day once, end at `finalThrough`.** `retrievableThrough` is
usually today, and today's row is the day so far. Recent days can still change
too, because the archive records days 2 to 3 behind live data. `finalThrough` is
the earlier of `recordedTo` and `retrievableThrough`: the newest day the archive
has recorded. Store through that, and fetch the days after it on your next run.
The archive can occasionally re-record a past day, for example after a park's
feed is repaired, so fetch a range again if you need to pick that up.

**`days()` yields, it does not collect.** Nothing accumulates, so the only thing
that grows is whatever you write the rows to.

Expand All @@ -332,6 +342,32 @@ try {
`history.changeRows(query)` is the same treatment for `changes`: one flattened
stream of `{ entityId, row }` whether you asked a park or a ride.

To rebuild what an entity was doing at any moment you also need the state before
its first change. That is `opening`, keyed by entity id, on the same result:

```js
const changes = history.changeRows({ date: '2026-09-20' });
for await (const { entityId, row } of changes) console.log(row.time, entityId, row.status);
for (const [entityId, opening] of Object.entries(changes.opening)) {
// In force from opening.time until this entity's first row.
console.log(entityId, opening.time, opening.status);
}
```

Each row holds from its `time` until the next row's, so `opening` plus the rows
cover the whole range with no gap. Without it, a ride still running from the
night before has no known status until its first change. There is an entry for
every entity in the response, including one that did not change all day, and
`opening.observedAt` says when that state was last seen, which can be long before
the range for a ride whose feed stopped. `opening` is readable once the response
has arrived: iterate first, or `await changes.load()`. Either way it is one
request. It is not enumerable, so spreading or logging the result is safe.

An opening with `degraded: true` is incomplete: the server could not look far
enough back for this response, and `degradedReason` says why (`timeout`,
`error` or `capacity`). A field it holds may be missing, so ask again in a
minute for the full opening before rebuilding a day from it.

### Or skip the code: there is a command

Installing the package puts `themeparks-backfill` on your path. It is the same
Expand All @@ -346,6 +382,7 @@ export THEMEPARKS_API_KEY=tpw_your_key
npx themeparks-backfill --list disney # find your park. This part needs no key.
npx themeparks-backfill "magic kingdom" # NDJSON, into the current directory
npx themeparks-backfill "Walt Disney World Resort" --format csv --out ./data
npx themeparks-backfill "magic kingdom" --since 2025-01-01 --until 2025-12-31
```

A park or a **destination**, by name or by id; a destination writes one file per
Expand All @@ -358,6 +395,35 @@ work it out. It checkpoints against the hourly history budget and exits 75
(`EX_TEMPFAIL`) when that runs out, so a cron or systemd timer retries instead of
alerting and the same command continues where it stopped. `--help` has the rest.

**Run it again to update.** A second run of a finished park fetches only the days
that have become final since the last one and appends them, so a nightly cron
keeps the file current. Only final days are written: a run ends at the newest day
the archive has finished recording and says so when it holds newer days back.
Each day is fetched once, as the archive recorded it.

**Fetching a range again.** The archive can occasionally re-record past days,
for example when a park's feed is repaired. A file never rewrites rows it
already holds, so to pick up a correction, download the affected days into a
separate directory and replace those `(entityId, date)` rows where you load the
data, or start the file again:

```bash
npx themeparks-backfill "Epcot" --since 2026-06-01 --until 2026-06-30 --out ./refetch
npx themeparks-backfill "Epcot" --overwrite # or: the whole file again
```

**Stopping a run at any point is safe.** The state file is written after every
page with the size of the file at that moment, and the next run first cuts off
anything written after it, so no day is ever appended twice. Ctrl-C and SIGTERM
exit 130 and 143; a killed run needs nothing either. Two runs on the same park
and `--out` at once are refused.

**`--since` and `--until`** (`YYYY-MM-DD`, both inclusive) pick the days, instead
of everything your plan reaches. `--since` applies when a file is started: a
later run accepts the same `--since` or a later one, so a fixed or a rolling cron
line both work, and refuses one earlier than the file's first day, or one that
would leave a gap, with what to do instead. `--overwrite` starts the file again.

The CSV is byte-for-byte identical to the Python SDK's, which runs the same
command: Magic Kingdom's five-year archive is 94,223 rows and 41 columns from
either.
Expand Down
15 changes: 11 additions & 4 deletions examples/backfill.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,12 @@
* resort is around a hundred times fewer calls than the same data pulled
* ride by ride.
*
* 2. It bounds the range with span().retrievableThrough, not with what the
* archive holds. Those are different dates on every plan below the top one,
* and asking past the entitlement is how a long backfill ends in 403s.
* 2. It ends at span().finalThrough: the earlier of what your key may read
* (retrievableThrough) and what the archive holds (recordedTo). Asking past
* the entitlement is how a long backfill ends in 403s, and the days past
* recordedTo are not final: today's row is the day so far, and the archive
* records days 2 to 3 behind live data. Stopping there means each day is
* written once, as the archive recorded it.
*
* 3. It checkpoints. The history budget is hourly, so a spent one can be most
* of an hour from resetting. The SDK raises BudgetExhaustedError rather
Expand Down Expand Up @@ -98,7 +101,11 @@ async function backfillPark(tp, parkId, outDir, format) {
const resuming = existsSync(checkpointPath);
const hasRows = existsSync(outPath) && statSync(outPath).size > 0;
const from = resuming ? readFileSync(checkpointPath, 'utf8').trim() : span.archiveFrom;
const to = span.retrievableThrough;
const to = span.finalThrough;
if (to === null) {
console.error(`${parkId}: nothing final to fetch yet; run again later`);
return 0;
}

console.error(`${parkId}: ${from} .. ${to}${resuming ? ' (resumed)' : ''} -> ${outPath}`);

Expand Down
24 changes: 17 additions & 7 deletions src/backfill-cli.ts
Original file line number Diff line number Diff line change
Expand Up @@ -17,14 +17,24 @@
*
* A file whose only job is to run has no condition to get wrong.
*/
import { run } from './backfill.js';
import { releaseLocks, run } from './backfill.js';

// Ctrl-C says how to continue, like the Python SDK. 130 is the shell's convention
// for SIGINT, and the state files on disk already say where each park got to.
process.on('SIGINT', () => {
process.stderr.write('\nstopped. Run the same command again to continue.\n');
process.exit(130);
});
// Ctrl-C, or a scheduler's SIGTERM, says how to continue, like the Python SDK.
// Exiting straight away is safe: the state file is rewritten atomically after
// every complete page, and the next run cuts off any rows written after that
// checkpoint before resuming, so an interrupted page is fetched again rather
// than appended twice. The locks are released so the next run need not wait
// to find them stale. 130 and 143 are the shell's conventions (128 + signal).
for (const [signal, code] of [
['SIGINT', 130],
['SIGTERM', 143],
] as const) {
process.on(signal, () => {
releaseLocks();
process.stderr.write('\nstopped. Run the same command again to continue.\n');
process.exit(code);
});
}

run().then(
(code) => {
Expand Down
Loading
Loading