Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
111 changes: 111 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,116 @@
# Changelog

## [Unreleased]

### Added

- **`themeparks-backfill --since YYYY-MM-DD` and `--until YYYY-MM-DD`.** A key
that reaches the whole archive used to download all of it, every time: a
buyer who wanted the last twelve months got five years. Both days are
inclusive and must be real calendar days written `YYYY-MM-DD`; `--since`
after `--until` is refused before anything is requested. `--since` applies when a file is started. A later run accepts the
same `--since` or a later one (a cron line computing "30 days ago" works), and
refuses one earlier than the file's first day, or one that would leave a gap,
instead of quietly handing back a file that is not what was asked for.

- **`history.changes()` exposes the `opening` state.** The raw history
response carries, per entity, the state in force at the start of the range,
and `changes()` threw it away. Without it the minutes between midnight and an
entity's first change had no known status, so a day rebuilt from raw history
disagreed with the daily summary whenever a ride was still running from the
night before. The result of `changes()` now has `opening`, a dict of
`HistoryOpening` keyed by entity id, covering every entity in the response,
including one that did not change all day. Iterating it yields exactly what it
always did, and reading `opening` costs no extra request. The async client's
result has the same attribute once the response has arrived: iterate first,
or `await changes.load()`.

- **`HistorySpan.final_through`**: the newest day the archive has recorded that
the key may read, the earlier of `recorded_to` and `retrievable_through`: the
place to stop if you fetch each day once. A property, so `a, b, c = span`
still works.

- **`HistoryOpening`** is exported, and `HistoryChanges.close()` /
`AsyncHistoryChanges.aclose()` end iteration early, as they did on the
generators these replaced.

- **An interrupted `themeparks-backfill` never appends a day twice.** The state
file is written before the first request and after every page, atomically,
with the size of the data file at that moment; the next run first cuts the
file back to that size. Ctrl-C, SIGTERM (now exit 143) and SIGKILL at any point
cost at most the page in flight. Two runs on the same park and `--out` at once
are refused (POSIX).

### Fixed

- **A finished park now updates on the next run.** A rerun used to print
`already complete` and exit 0 without fetching a single new day, so a nightly
cron looked healthy and never updated; the only way to get yesterday was
`--overwrite`, which downloaded the whole archive again. A finished file is now
carried forward from the day after its last one, appending only the new days.
An interrupted run still resumes exactly where it stopped, including a rerun
interrupted before its first page, which would otherwise have started again
from the top of the archive.

- **The newest rows of a backfill were partial days, and stayed that way.** A run
ended on `retrievableThrough`, which is usually today: today's row is the day
so far, and the archive records days 2 to 3 behind live data, so the last few
days of every file were still changing when they were written. Magic Kingdom's
last day summed to about half the operating minutes of a full one. A run now
ends at `final_through`, says so when it holds days back, and the next run adds
them once they are final. Each day is fetched once, as the archive recorded
it; the README says how to fetch a range again if the archive later
re-records it.

**Files written by 4.0.x are corrected once.** Their state file does not say
which of their newest days were final, so the first run of this version
removes the rows from the last seven days before that run's end and fetches
those days again. Every other row is left byte for byte as it was. A 4.0 file
whose newest row is older than that is not rewritten at all, and one that lies
wholly inside those seven days, as every anonymous 7-day file does, is simply
downloaded again. The state file format moves to version 2 for this; version 1
files from this SDK are upgraded, not refused.

- **An interrupted nightly extension appended the same days twice.** The state
was written only at the end of a run or on an SDK error, so Ctrl-C, SIGTERM
or SIGKILL during an extension left it saying finished through the old day,
and the rerun appended those days again. See the checkpointing above.

- **A run that stopped part-way through its first page lost days.** It resumed
from the newest day any entity had reached; rows arrive entity by entity, so
the entities behind it lost the days in between. Checkpoints make this
impossible for new files, and an older state with only that day goes back a
whole page (31 days) instead.

- **The state recorded the start asked for, not the first day written.** After
the key's window moved it later, a `--since` earlier than the real first day
passed silently. `start` is now the first day written and a new `since`
field keeps the one asked for, so the same `--since` keeps working and a
different one before the file is refused.

- **A continued file could skip ahead to the key's first day.** When the day a
file continues from is older than the key may read, because a cron missed more
days than the window or a plan lapsed, the run carried on from the key's first
day and left a gap the state file did not record. It is refused now with exit
1, the file and its state untouched, and a message naming `--overwrite`. A
fixed `--since` older than the window is not affected: the file starts at the
key's first day, and the same command line keeps working every night.

- **A finished file written to a different column layout was appended to.** Only
an unfinished one was refused. A finished one fell through to a fresh start,
which opened the existing file in append mode and wrote the whole archive into
it a second time under a second header, exit 0. It is refused now, the same
way.

- **A state file whose data file had been deleted was continued**, producing a
file that started part-way through its range and was then recorded as
complete. The park is downloaded again from the start instead.

- **The README said `waitTime` is an `int`.** The API's schema declares it a
JSON `number`, the models type it `float`, and a raw row dumped to JSON says
`45.0`. The README now says `float | None`, explains the `45.0`, and a test
holds its table to the models' types.

## [4.0.1] - 2026-09-28

### Fixed
Expand Down
80 changes: 73 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,13 +147,19 @@ Each variant is exposed as an attribute on `entry.queue`. All are

| Attribute | Type | Fields |
|------------------------|-----------------------|------------------------------------------------------------------------------------------------------|
| `queue.STANDBY` | `StandbyQueue` | `waitTime: int \| None` |
| `queue.SINGLE_RIDER` | `SingleRiderQueue` | `waitTime: int \| None` |
| `queue.PAID_STANDBY` | `PaidStandbyQueue` | `waitTime: int \| None` |
| `queue.STANDBY` | `StandbyQueue` | `waitTime: float \| None` |
| `queue.SINGLE_RIDER` | `SingleRiderQueue` | `waitTime: float \| None` |
| `queue.PAID_STANDBY` | `PaidStandbyQueue` | `waitTime: float \| None` |
| `queue.RETURN_TIME` | `ReturnTimeQueue` | `state`, `returnStart`, `returnEnd` |
| `queue.PAID_RETURN_TIME` | `PaidReturnTimeQueue` | `state`, `returnStart`, `returnEnd`, `price` |
| `queue.BOARDING_GROUP` | `BoardingGroupQueue` | `allocationStatus`, `currentGroupStart`, `currentGroupEnd`, `nextAllocationTime`, `estimatedWait` |

`waitTime` is a `float` because the API's schema declares it a JSON `number`,
not an integer, so a 45-minute wait arrives as `45.0` and a raw row dumped to
JSON says `45.0`. Use `current_wait_time(entry)` or `int(...)` when you want an
`int`, and format with `{wait:.0f}`, not `{wait:d}`, if you print the field
directly.

#### Direct access

```python
Expand Down Expand Up @@ -201,7 +207,7 @@ with ThemeParks() as tp:
live = tp.entity("75ea578a-adc8-4116-a54d-dccb60765ef9").live()
for entry in live.liveData or []:
for q in iter_queues(entry):
# q is a dict, e.g. {"type": "STANDBY", "waitTime": 35}
# q is a dict, e.g. {"type": "STANDBY", "waitTime": 35.0}
# or {"type": "PAID_RETURN_TIME", "state": "AVAILABLE", ...}
print(entry.name, q)
```
Expand Down Expand Up @@ -274,19 +280,32 @@ with ThemeParks(api_key=KEY) as tp:
history = tp.entity(DISNEYLAND).history

# What exists, and what your key may read. Same three fields whether the
# id is a park or a single ride.
# id is a park or a single ride. `final_through` is the newest day whose
# the archive has recorded: store through that, ask for the rest later.
span = history.span()
print(span.archive_from, span.recorded_to, span.retrievable_through)
print(span.final_through)

# One summary row per park-local day, as (entity id, row).
for entity_id, row in history.days(span.archive_from, span.retrievable_through):
print(row.date, entity_id, row.operatingMinutes, row.standby.p50 if row.standby else None)

# Every recorded change on one day.
for entity_id, row in history.changes("2026-09-20"):
# Every recorded change on one day, and the state before the first of them.
changes = history.changes("2026-09-20")
for entity_id, row in changes:
print(row.time, entity_id, row.status)
for entity_id, opening in changes.opening.items():
print(entity_id, "at the start of the day:", opening.status)
```

**A day rebuilds from `opening` plus the rows.** Each row is the entity's
complete state from its `time` until the next row. `changes.opening` is the
state in force before the first row, keyed by entity id, so the minutes between
midnight and a ride's first change have a status too: a ride still running from
the night before, say. It covers every entity in the response, including one
that did not change all day. Iterating `changes()` yields exactly what it
always did, and reading `opening` costs no extra request.

**Ask the park, not the rides.** Both history endpoints answer every entity in
a park in one request. Pulling the same data ride by ride is around a hundred
times more calls for a large resort, against the same budget. Pass a park id
Expand Down Expand Up @@ -331,6 +350,53 @@ It reads how far back your own key may ask and starts there, writes NDJSON or
has done so re-running never duplicates a file, and exits 75 when the hourly
history budget runs out so a scheduler retries rather than alerts.

```bash
themeparks-backfill "Epcot" --since 2025-01-01 # not the whole archive
themeparks-backfill "Epcot" --since 2025-01-01 --until 2025-12-31 # one year, both days inclusive
```

**Run it again to bring a file up to date.** A finished park is carried
forward from the day after its last one, so the same command in a nightly cron
appends the new days and nothing else. Only **final** days are written: today's
row is the day so far, and the archive records days 2 to 3 behind live data, so
the newest days can still change. A run stops at the newest final day
(`span().final_through`) and says so, and the next run adds the rest. Each day
is fetched once, as the archive recorded it.

**Fetching a range again.** The archive can occasionally re-record past days,
for example when a park's feed is repaired. A file never rewrites rows it
already holds, so to pick up a correction, download the affected days into a
separate directory and replace those `(entityId, date)` rows where you load the
data, or start the file again:

```bash
themeparks-backfill "Epcot" --since 2026-06-01 --until 2026-06-30 --out ./refetch
themeparks-backfill "Epcot" --overwrite # or: the whole file again
```

Stopping a run at any point is safe. The state file is written after every page
with the size of the file at that moment, and the next run first cuts off
anything written after it, so no day is ever appended twice. Ctrl-C and SIGTERM
exit 130 and 143; a killed run needs nothing either. Two runs on the same park
and `--out` at once are refused.

`--since` applies when a file is started. Later runs continue that file and
accept the same `--since`, or a later one, such as a cron line computing "30
days ago". One earlier than the file's first day, or one that would leave a gap,
is refused rather than ignored: pass `--overwrite`, or a different `--out`. A
fixed `--since` older than your key's window starts the file at the first day
your key can read, and the same command line keeps working every night.

A file is never continued past a gap. If the day it would continue from is
older than your key can read, because a cron missed more days than your window
or a plan lapsed, the run is refused with exit 1 and the file is left alone.

Files written by 4.0.x ended on today, so their newest rows can be partial. The
first run of this version removes the rows from the last week of such a file and
fetches those days again, final this time. Everything else in the file is left
exactly as it was. A 4.0 file that lies wholly inside that week, as every
anonymous 7-day file does, is simply downloaded again.

`python -m themeparks.backfill` is the same thing, which is the one to use if
`pip install --user` put the script somewhere off your PATH. `themeparks-backfill
--help` has the rest.
Expand Down
16 changes: 16 additions & 0 deletions docs/api/history.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,22 @@ hundred times fewer calls than the same data fetched ride by ride.
options:
heading_level: 2

::: themeparks.HistoryChanges
options:
heading_level: 2

::: themeparks.AsyncHistoryChanges
options:
heading_level: 2

::: themeparks.HistoryOpening
options:
heading_level: 2

::: themeparks.HistorySpan
options:
heading_level: 2

::: themeparks.BudgetExhaustedError
options:
heading_level: 2
38 changes: 34 additions & 4 deletions docs/cookbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -283,7 +283,7 @@ with ThemeParks() as tp:
live = tp.entity(MAGIC_KINGDOM).live()
for entry in live.liveData or []:
for q in iter_queues(entry):
# q is e.g. {"type": "STANDBY", "waitTime": 45}
# q is e.g. {"type": "STANDBY", "waitTime": 45.0}
# or {"type": "PAID_RETURN_TIME", "state": "AVAILABLE", "price": {...}, ...}
print(entry.name, q["type"], q)
```
Expand All @@ -306,7 +306,10 @@ last day your key may retrieve, so you ask for days that exist instead of
discovering the ends by trial. Bound the backfill by `retrievable_through`,
not by `recorded_to`: the archive holds more than a free or Pro key is
entitled to read, and asking past the entitlement is how a backfill walks into
a wall of 403s at the end of a long run.
a wall of 403s at the end of a long run. If you write each day once and never
revisit it, end at `span.final_through` instead: the earlier of the two, and
the newest day the archive has recorded. The archive can re-record a past day
after a feed repair, so fetch a range again if you need to pick that up.

```python
import json
Expand Down Expand Up @@ -392,8 +395,15 @@ range above is refused on its FIRST request unless you already know your floor.
The command reads it out of the 403 and starts again there.

**Re-running.** The loop above appends, so running it twice doubles the file.
The command records what it wrote and declines to fetch a finished park again
unless you pass `--overwrite`.
The command records what it wrote, and a second run carries a finished park
forward from the day after its last one instead, so a nightly cron keeps the
file current. `--since` and `--until` pick the days; `--overwrite` starts again.

**Days that are not final yet.** The loop above ends at `retrievable_through`,
usually today, and today's row is the day so far. Recent days can still change
too, because the archive records days 2 to 3 behind live data. Stop at
`span.final_through` if you store rows once, as the command does, or fetch the
days after it again on your next run.

It also takes destinations as well as parks, names every row with the park and
the entity as the history response reported them, and has `--list` for finding
Expand All @@ -413,6 +423,26 @@ with ThemeParks(api_key="YOUR_KEY") as tp:
print(row.time, entity_id, row.status, row.queue)
```

To rebuild what an entity was doing at any moment, you also need the state
before its first change. That is `opening`, keyed by entity id, on the same
result:

```python
with ThemeParks(api_key="YOUR_KEY") as tp:
changes = tp.entity(DISNEYLAND).history.changes("2026-09-20")
rows = list(changes)
for entity_id, opening in changes.opening.items():
# In force from opening.time until this entity's first row.
print(entity_id, opening.time, opening.status)
```

Each row holds from its `time` until the next row's, so `opening` plus the
rows cover the whole range with no gap. Without it, a ride that was still
running from the night before has no known status until its first change.
`opening.observedAt` says when that state was last seen, which can be long
before the range for a ride whose feed stopped. With the async client, iterate
first or call `await changes.load()` before reading `opening`.

A park answers one day per call. A single entity answers up to 31 days, so
pass `start=` and `end=` there instead of `date=`. You do not have to
remember which cap applies: ask for the range you want, and the API either
Expand Down
14 changes: 14 additions & 0 deletions tests/fixtures/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,3 +48,17 @@ the server says carry on at 2026-09-01, but two of its three entities have no
rows after 2026-08-30 -- so a checkpoint taken from the newest ROW rewinds and
re-downloads days already written. The same files are in the JavaScript SDK, so
both ports are tested against identical bytes.

## space_mountain_history_2026-09-26.json / space_mountain_daily_2026-09-26.json

`GET /entity/b2260923-9315-40fd-9c6b-44dd811dbe64/history?date=2026-09-26` and
`GET /entity/b2260923-9315-40fd-9c6b-44dd811dbe64/history/daily?date=2026-09-26`,
both captured 2026-09-28 without a key, verbatim.

They are the oracle for `changes().opening`. Space Mountain's opening that day is
`OPERATING`, because the previous night's hours ran past midnight, and its first
row is the close at 00:01:03 local. A day rebuilt from the rows alone has 63
seconds with no known status; rebuilt from the opening plus the rows it has
none, and its first open and last close are the daily row's `firstOperatingAt`
and `lastClosedAt`. If a re-capture picks a day whose opening is `CLOSED`, the
tests stop being able to tell the two apart, and one of them says so.
Loading
Loading