diff --git a/CHANGELOG.md b/CHANGELOG.md index 2b269f6..aa80013 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,138 @@ # Changelog +## [4.0.0] - 2026-09-28 + +**3.3.0 was yanked: incomplete CSV export and a resume defect.** + +A major version because **the CSV header changed**: fifteen columns were added and +the order is now the schema's, so a reader that takes columns by position gets the +wrong ones rather than an error. Read by name. The library API is backward +compatible. + +Everything here came out of porting `themeparks-backfill` to the JavaScript SDK +and then diffing the two outputs over the same park, and out of six reviews of the +result. Two independent implementations reading one API disagree in exactly the +places one of them is wrong. Magic Kingdom's full archive now comes back +**byte for byte identical** from both SDKs: 94,223 rows, 41 columns, the only +differences being today's row, which grows as the day elapses. + +### Fixed + +- **`themeparks-backfill "magic kingdom"` wrote the wrong park name into every + row.** A name that matched one park by substring returned the formatted display + label, so `parkName` read `Magic Kingdom Park (Walt Disney World® Resort)` for + all ~94,000 rows, and the resolution echo printed the destination twice. Four + live names reached it. + +- **The CSV was missing ten of the thirty-six fields the API sends, on every row.** + `unknownMinutes`, the whole `inParkHours` block (the day's numbers limited to the + park's published hours -- usually the ones you want, since a ride "down" at 2am + is not down), `extremeWaits` (how many readings of 480+ minutes are folded into + the statistics, which is how you spot a feed error), and three of `singleRider`'s + five percentiles while `standby` carried all five. On a five-year Magic Kingdom + export, 72,200 of 94,223 rows were missing their in-park statistics. **The column + list is now derived from the model**, so it cannot drift again. + +- **Vendored models were stale, and pydantic drops what it does not declare**, so + those three fields were deleted at parse time for every caller of `days()`, not + just for the CSV. Models regenerated, and every model now keeps fields the schema + does not declare (`themeparks._models_base.ApiModel`, `extra="allow"`), so a + field the API adds tomorrow survives parsing and reaches `model_dump()` and the + NDJSON output before this SDK knows it exists. It does not reach the CSV, whose + columns come from the schema. + +- **A resumed download duplicated a day.** The checkpoint was the newest row + written; the page it came from covered further, because an entity that stopped + reporting has no rows for the tail days. A rerun re-fetched a day already in the + file and appended every row of it again, breaking the `(entityId, date)` key -- + on the exit-75 path, which is the ordinary path for a long back fill. The + checkpoint is now the day the server's own `next` URL starts on. + +- **A failure on a resumed run deleted everything already downloaded.** `written + == 0` means "this process wrote nothing", not "the file is empty". The state file + survived pointing mid-archive, so the next run appended only the tail and + recorded `complete: true`. Same for a window that closes under a resumed run -- + a key rotated out of a scheduler's environment, a lapsed subscription -- which + additionally exited 0, so the scheduler logged success, and became a permanent + trap. + +- **Resuming across versions, formats or SDKs corrupted the file.** One state file + served both formats, so `ndjson` then `csv` then `ndjson` doubled every row in + the first file; and the state carried nothing about the header, so 3.3.0's + 19-column file resumed under this build appended 41-field rows beneath it. The + state file is now `..backfill-state.json` and records the SDK, + its version, a state version and a fingerprint of the exact header. Anything that + does not match is refused with a message saying why, never resumed. + +- **A network failure or timeout now exits 75, not 1**, so a scheduler retries + rather than alerting; anything the API actively rejected still exits 1. The + JavaScript SDK had these the other way round. + +- **A carriage return in an entity name was written unquoted on Python 3.9 and + 3.10**, so one row parsed as two with every later column shifted. The `csv` + module's QUOTE_MINIMAL only quotes characters that appear in the line terminator, + and this command sets LF; 3.11 changed the module to always quote CR and LF, so + the defect was invisible on a modern interpreter and live on two supported ones. + The CSV writer now does its own minimal quoting, which also makes the output + byte-identical across Python versions rather than only within one. + +- **UTC timestamps are written `Z`, not `+00:00`**, and CSV line endings are LF. + Between them these accounted for 39,201 differing lines against the JavaScript + SDK's output for no difference in meaning. + +- **One park's failure no longer abandons the rest of a destination.** Every park + is tried, what failed is named at the end, and the exit code still says something + went wrong. A spent budget still stops everything, deliberately. + +- **A failed park no longer leaves a 0-byte file** that reads as "this park has no + history", including when the budget runs out before the first page. + +- **A network failure, a full disk or Ctrl-C is a sentence, not a traceback.** + +- **The user agent named neither version.** It was the literal + `themeparks-backfill/1`, and it replaced the SDK's own, so a support question had + no version to work from at either end. + +- **`--list ` reported the wrong total**, printing "all 1 parks" for a + destination with six -- on the one line whose whole job is that number. + +- **An ambiguous name listed the wrong candidates**, widening to substrings and + offering a third park that was not what was typed. It now lists the ids of the + parks that actually match, sorted by name. + +- **A collection of nested models would have produced phantom columns** and then an + `AttributeError` on the first row. Duplicate column names are now impossible at + import rather than a wrong number under a right-looking header. + +- **The NDJSON identity columns could be overwritten by the row** once models kept + undeclared fields. + +### Added + +- **The CSV carries a UTF-8 BOM**, so Excel on Windows stops rendering + `Walt Disney World® Resort` as mojibake. +- **A cell a spreadsheet would execute is prefixed with an apostrophe** (`=`, `+`, + `-`, `@`, tab, CR). Numeric cells are left alone, so a negative number stays a + number. +- **`on_page` on `days()` and `days_with_entities()`**, called once every row of a + page has been yielded, with a `HistoryPage` (`start`, `end`, `next_url`). The page + boundary is the server's own answer to "where do I carry on", and the rows cannot + tell you. +- **`--version`.** +- `EntityRef` and `HistoryPage` are exported from the package. +- `tests/fixtures/csv_contract.json`, an identical copy of which lives in the + JavaScript SDK. Both suites assert their column list against it, because this is + one command with two implementations and a customer using both should get one + file format. + +### Changed + +- The `themeparks-backfill` entry point is `themeparks.backfill:cli`, which adds + the top-level error handling. `main()` is unchanged for anyone calling it. +- Model equality and `model_json_schema()` reflect `extra="allow"`: two responses + differing only in an undeclared field now compare unequal, and dumps may contain + fields the schema does not list. + ## [3.3.0] - 2026-09-28 ### Added diff --git a/pyproject.toml b/pyproject.toml index e0d4f2b..1a42116 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "themeparks" -version = "3.3.0" +version = "4.0.0" description = "Official SDK for the ThemeParks.wiki API" readme = "README.md" requires-python = ">=3.9" @@ -34,7 +34,7 @@ dependencies = [ # The archive backfill, as a command rather than a file to copy off GitHub. # `pip install themeparks` then `themeparks-backfill "Disneyland Park"` is the # whole path from nothing to a file of history. -themeparks-backfill = "themeparks.backfill:main" +themeparks-backfill = "themeparks.backfill:cli" [project.urls] Homepage = "https://api.themeparks.wiki" diff --git a/scripts/regenerate.py b/scripts/regenerate.py index 5d921b2..76fcafd 100644 --- a/scripts/regenerate.py +++ b/scripts/regenerate.py @@ -136,7 +136,12 @@ def apply_nullable_patches(text: str) -> str: # result was a model that cannot parse the last page of any paged # response, where `next` is null. pattern = re.compile( - rf"(?Pclass {class_name}\(BaseModel\):\n" + # The base class is `ApiModel`, ours, not `BaseModel` -- see + # themeparks/_models_base.py for why. Both are accepted so the patch + # does not silently stop matching if that changes again; the + # invariant check below is what turns a miss into a hard stop, and it + # did exactly that when the base class moved. + rf"(?Pclass {class_name}\((?:BaseModel|ApiModel)\):\n" rf"(?:(?: [^\n]*)?\n)*?" rf" {field_name}:\s*)" rf"(?P[^\n]+)" @@ -182,6 +187,12 @@ def main() -> None: str(OUTPUT), "--output-model-type", "pydantic_v2.BaseModel", + # EVERY MODEL KEEPS WHAT THE SPEC DOES NOT DECLARE. Pydantic's default is + # to drop it, and this spec trails the API by days at a time, so during + # that window new fields were being deleted at parse time for every + # caller -- silently, with nothing failing. See themeparks/_models_base.py. + "--base-class", + "themeparks._models_base.ApiModel", "--use-schema-description", "--use-field-description", "--use-annotated", @@ -216,6 +227,19 @@ def main() -> None: # Auto-format the generated file so `ruff format --check` in CI doesn't # fail on quote-style or whitespace differences from datamodel-codegen. print("Formatting generated models with ruff...") + # `check --fix` first, for the import ordering: the custom base class means + # the generator emits a first-party import in among the third-party ones, and + # `format` alone does not sort imports. Without this, `ruff check` in CI fails + # on a file nobody is allowed to edit by hand. + subprocess.run( + # --select I ONLY. Unpinned, this is one `pyproject.toml` edit away from + # deleting imports: the F401 per-file-ignore for `_generated/*` is the only + # reason a stray `import os` survives it today, and this step runs with + # check=False inside the nightly drift job, so a tightened ignore would + # start removing code silently. + [sys.executable, "-m", "ruff", "check", "--select", "I", "--fix", "--quiet", str(OUTPUT)], + check=False, + ) subprocess.run( [sys.executable, "-m", "ruff", "format", str(OUTPUT)], check=True, diff --git a/tests/fixtures/README.md b/tests/fixtures/README.md index a19940c..253ebc4 100644 --- a/tests/fixtures/README.md +++ b/tests/fixtures/README.md @@ -35,3 +35,16 @@ It is trimmed to keep, deliberately, every case that broke name resolution: **Do not edit these by hand.** Re-capture them. If a name upstream has drifted, that is a real change and the test should notice. + +## mk_park_daily_page1.json / mk_park_daily_page2.json + +Two consecutive pages of one real request, captured 2026-09-28: +`GET /entity/75ea578a-adc8-4116-a54d-dccb60765ef9/history/daily?from=2026-08-01&to=2026-09-20` +then its `next` followed verbatim. Trimmed to three entities (an attraction, a +show, a restaurant); `range`, `next` and every row are the server's. + +They are the oracle for resumable paging. Page one covers through 2026-08-31 and +the server says carry on at 2026-09-01, but two of its three entities have no +rows after 2026-08-30 -- so a checkpoint taken from the newest ROW rewinds and +re-downloads days already written. The same files are in the JavaScript SDK, so +both ports are tested against identical bytes. diff --git a/tests/fixtures/csv_contract.json b/tests/fixtures/csv_contract.json new file mode 100644 index 0000000..8c5263b --- /dev/null +++ b/tests/fixtures/csv_contract.json @@ -0,0 +1,58 @@ +{ + "_comment": [ + "THE CSV CONTRACT, shared by the Python and JavaScript SDKs.", + "Both repos hold an identical copy of this file and assert their own column", + "list against it, because `themeparks-backfill` is one command with two", + "implementations and a customer using both must get one file format.", + "Before this existed, one SDK wrote 32 columns and the other 41, with", + "inParkScheduledMinutes against inParkHoursScheduledMinutes, and the test that", + "claimed to check it was four spot-checks.", + "Generated from the OpenAPI spec: run `npm run regenerate` in the JavaScript", + "SDK, copy this file to both repos, and expect both suites to fail until they", + "agree." + ], + "fingerprint": "1b6ea478049dc7ca", + "columns": [ + "parkId", + "parkName", + "entityId", + "entityName", + "entityType", + "date", + "firstOperatingAt", + "lastClosedAt", + "operatingMinutes", + "downMinutes", + "unknownMinutes", + "standbyMin", + "standbyP50", + "standbyMean", + "standbyP90", + "standbyMax", + "singleRiderMin", + "singleRiderP50", + "singleRiderMean", + "singleRiderP90", + "singleRiderMax", + "extremeWaitsStandby", + "extremeWaitsSingleRider", + "showCount", + "inParkHoursScheduledMinutes", + "inParkHoursOperatingMinutes", + "inParkHoursDownMinutes", + "inParkHoursUnknownMinutes", + "inParkHoursStandbyMin", + "inParkHoursStandbyP50", + "inParkHoursStandbyMean", + "inParkHoursStandbyP90", + "inParkHoursStandbyMax", + "inParkHoursSingleRiderMin", + "inParkHoursSingleRiderP50", + "inParkHoursSingleRiderMean", + "inParkHoursSingleRiderP90", + "inParkHoursSingleRiderMax", + "inParkHoursExtremeWaitsStandby", + "inParkHoursExtremeWaitsSingleRider", + "changes" + ] +} diff --git a/tests/fixtures/mk_park_daily_page1.json b/tests/fixtures/mk_park_daily_page1.json new file mode 100644 index 0000000..3aa3c9b --- /dev/null +++ b/tests/fixtures/mk_park_daily_page1.json @@ -0,0 +1,1465 @@ +{ + "id": "75ea578a-adc8-4116-a54d-dccb60765ef9", + "name": "Magic Kingdom Park", + "entityType": "PARK", + "parentId": "e957da41-3552-4cf6-b636-5babc5cbc4e5", + "destinationId": "e957da41-3552-4cf6-b636-5babc5cbc4e5", + "timezone": "America/New_York", + "range": { + "from": "2026-08-01", + "to": "2026-08-31" + }, + "entities": [ + { + "id": "f5aad2d4-a419-4384-bd9a-42f86385c750", + "name": "\"it's a small world\"", + "entityType": "ATTRACTION", + "coverage": { + "firstRecordedAt": "2021-07-03" + }, + "days": [ + { + "date": "2026-08-01", + "firstOperatingAt": "2026-08-01T12:31:02Z", + "lastClosedAt": "2026-08-02T03:00:14Z", + "operatingMinutes": 734, + "downMinutes": 134, + "unknownMinutes": 1, + "standby": { + "min": 5, + "p50": 10, + "mean": 14, + "p90": 30, + "max": 40 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 734, + "downMinutes": 134, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 14, + "p90": 30, + "max": 40 + } + }, + "changes": 177 + }, + { + "date": "2026-08-02", + "firstOperatingAt": "2026-08-02T12:31:10Z", + "lastClosedAt": "2026-08-03T03:00:16Z", + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 20, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 20, + "max": 30 + } + }, + "changes": 191 + }, + { + "date": "2026-08-03", + "firstOperatingAt": "2026-08-03T12:31:26Z", + "lastClosedAt": "2026-08-04T03:00:24Z", + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 10, + "mean": 14, + "p90": 25, + "max": 35 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 14, + "p90": 25, + "max": 35 + } + }, + "changes": 191 + }, + { + "date": "2026-08-04", + "firstOperatingAt": "2026-08-04T12:31:30Z", + "lastClosedAt": "2026-08-05T03:00:34Z", + "operatingMinutes": 859, + "downMinutes": 9, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 15, + "mean": 16, + "p90": 30, + "max": 35 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 859, + "downMinutes": 9, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 16, + "p90": 30, + "max": 35 + } + }, + "changes": 203 + }, + { + "date": "2026-08-05", + "firstOperatingAt": "2026-08-05T12:30:35Z", + "lastClosedAt": null, + "operatingMinutes": 882, + "downMinutes": 47, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 17, + "p90": 30, + "max": 40 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 882, + "downMinutes": 47, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 17, + "p90": 30, + "max": 40 + } + }, + "changes": 188 + }, + { + "date": "2026-08-06", + "firstOperatingAt": "2026-08-06T12:31:45Z", + "lastClosedAt": "2026-08-07T03:00:52Z", + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 10, + "mean": 18, + "p90": 45, + "max": 50 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 18, + "p90": 45, + "max": 50 + } + }, + "changes": 187 + }, + { + "date": "2026-08-07", + "firstOperatingAt": "2026-08-07T16:01:57Z", + "lastClosedAt": null, + "operatingMinutes": 656, + "downMinutes": 271, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 5, + "mean": 11, + "p90": 25, + "max": 35 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 656, + "downMinutes": 271, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 11, + "p90": 25, + "max": 35 + } + }, + "changes": 69 + }, + { + "date": "2026-08-08", + "firstOperatingAt": "2026-08-08T11:31:03Z", + "lastClosedAt": "2026-08-09T03:00:11Z", + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 15, + "mean": 14, + "p90": 25, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 14, + "p90": 25, + "max": 30 + } + }, + "changes": 211 + }, + { + "date": "2026-08-09", + "firstOperatingAt": "2026-08-09T12:31:17Z", + "lastClosedAt": "2026-08-10T03:00:27Z", + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 10, + "p90": 20, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 10, + "p90": 20, + "max": 30 + } + }, + "changes": 202 + }, + { + "date": "2026-08-10", + "firstOperatingAt": "2026-08-10T12:31:24Z", + "lastClosedAt": "2026-08-11T03:00:31Z", + "operatingMinutes": 861, + "downMinutes": 7, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 10, + "p90": 20, + "max": 40 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 861, + "downMinutes": 7, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 10, + "p90": 20, + "max": 40 + } + }, + "changes": 212 + }, + { + "date": "2026-08-11", + "firstOperatingAt": "2026-08-11T11:32:08Z", + "lastClosedAt": null, + "operatingMinutes": 986, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 927, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 161 + }, + { + "date": "2026-08-12", + "firstOperatingAt": "2026-08-12T12:31:25Z", + "lastClosedAt": null, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 20, + "mean": 17, + "p90": 30, + "max": 40 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 20, + "mean": 17, + "p90": 30, + "max": 40 + } + }, + "changes": 200 + }, + { + "date": "2026-08-13", + "firstOperatingAt": "2026-08-13T12:30:33Z", + "lastClosedAt": "2026-08-14T03:00:41Z", + "operatingMinutes": 929, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 5, + "p50": 10, + "mean": 12, + "p90": 15, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 12, + "p90": 15, + "max": 30 + } + }, + "changes": 180 + }, + { + "date": "2026-08-14", + "firstOperatingAt": "2026-08-14T11:30:49Z", + "lastClosedAt": null, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 5, + "p50": 5, + "mean": 8, + "p90": 10, + "max": 20 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 5, + "mean": 8, + "p90": 10, + "max": 20 + } + }, + "changes": 155 + }, + { + "date": "2026-08-15", + "firstOperatingAt": "2026-08-15T11:30:56Z", + "lastClosedAt": "2026-08-16T03:01:01Z", + "operatingMinutes": 737, + "downMinutes": 192, + "unknownMinutes": 3, + "standby": { + "min": 5, + "p50": 15, + "mean": 13, + "p90": 25, + "max": 35 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 737, + "downMinutes": 192, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 15, + "mean": 13, + "p90": 25, + "max": 35 + } + }, + "changes": 164 + }, + { + "date": "2026-08-16", + "firstOperatingAt": "2026-08-16T12:31:12Z", + "lastClosedAt": "2026-08-17T02:00:15Z", + "operatingMinutes": 808, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 5, + "p50": 10, + "mean": 9, + "p90": 15, + "max": 20 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 808, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 9, + "p90": 15, + "max": 20 + } + }, + "changes": 179 + }, + { + "date": "2026-08-17", + "firstOperatingAt": "2026-08-17T12:30:28Z", + "lastClosedAt": "2026-08-18T03:00:25Z", + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 15, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 15, + "max": 25 + } + }, + "changes": 193 + }, + { + "date": "2026-08-18", + "firstOperatingAt": "2026-08-18T11:31:25Z", + "lastClosedAt": null, + "operatingMinutes": 907, + "downMinutes": 50, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 10, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 878, + "downMinutes": 50, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 15, + "max": 25 + } + }, + "changes": 157 + }, + { + "date": "2026-08-19", + "firstOperatingAt": "2026-08-19T12:31:42Z", + "lastClosedAt": null, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 5, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 30 + } + }, + "changes": 196 + }, + { + "date": "2026-08-20", + "firstOperatingAt": "2026-08-20T12:30:52Z", + "lastClosedAt": "2026-08-21T03:00:55Z", + "operatingMinutes": 929, + "downMinutes": 0, + "unknownMinutes": 3, + "standby": { + "min": 0, + "p50": 20, + "mean": 18, + "p90": 35, + "max": 45 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 20, + "mean": 18, + "p90": 35, + "max": 45 + } + }, + "changes": 202 + }, + { + "date": "2026-08-21", + "firstOperatingAt": "2026-08-21T11:30:58Z", + "lastClosedAt": null, + "operatingMinutes": 957, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 929, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 145 + }, + { + "date": "2026-08-22", + "firstOperatingAt": "2026-08-22T11:31:09Z", + "lastClosedAt": "2026-08-23T03:00:21Z", + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 3, + "standby": { + "min": 5, + "p50": 10, + "mean": 11, + "p90": 20, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 11, + "p90": 20, + "max": 25 + } + }, + "changes": 206 + }, + { + "date": "2026-08-23", + "firstOperatingAt": "2026-08-23T11:30:23Z", + "lastClosedAt": null, + "operatingMinutes": 958, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 7, + "p90": 10, + "max": 20 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 929, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 7, + "p90": 15, + "max": 20 + } + }, + "changes": 159 + }, + { + "date": "2026-08-24", + "firstOperatingAt": "2026-08-24T12:31:33Z", + "lastClosedAt": "2026-08-25T03:27:44Z", + "operatingMinutes": 836, + "downMinutes": 58, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 10, + "mean": 12, + "p90": 25, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 810, + "downMinutes": 58, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 30 + } + }, + "changes": 204 + }, + { + "date": "2026-08-25", + "firstOperatingAt": "2026-08-25T12:30:46Z", + "lastClosedAt": null, + "operatingMinutes": 898, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 10, + "max": 15 + } + }, + "changes": 136 + }, + { + "date": "2026-08-26", + "firstOperatingAt": "2026-08-26T12:32:07Z", + "lastClosedAt": null, + "operatingMinutes": 752, + "downMinutes": 115, + "unknownMinutes": 61, + "standby": { + "min": 0, + "p50": 10, + "mean": 11, + "p90": 20, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 752, + "downMinutes": 115, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 11, + "p90": 20, + "max": 30 + } + }, + "changes": 135 + }, + { + "date": "2026-08-27", + "firstOperatingAt": "2026-08-27T13:01:18Z", + "lastClosedAt": null, + "operatingMinutes": 803, + "downMinutes": 7, + "unknownMinutes": 511, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 25, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 803, + "downMinutes": 7, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 25, + "max": 25 + } + }, + "changes": 177 + }, + { + "date": "2026-08-28", + "firstOperatingAt": "2026-08-28T12:30:32Z", + "lastClosedAt": null, + "operatingMinutes": 898, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 10, + "max": 10 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 10 + } + }, + "changes": 136 + }, + { + "date": "2026-08-29", + "firstOperatingAt": "2026-08-29T12:30:36Z", + "lastClosedAt": "2026-08-30T03:00:43Z", + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 20, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 10, + "p90": 20, + "max": 25 + } + }, + "changes": 224 + }, + { + "date": "2026-08-30", + "firstOperatingAt": "2026-08-30T12:30:52Z", + "lastClosedAt": null, + "operatingMinutes": 898, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 10 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 10 + } + }, + "changes": 142 + }, + { + "date": "2026-08-31", + "firstOperatingAt": "2026-08-31T12:31:00Z", + "lastClosedAt": "2026-09-01T02:01:13Z", + "operatingMinutes": 809, + "downMinutes": 0, + "unknownMinutes": 3, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 10 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 809, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 10 + } + }, + "changes": 196 + } + ] + }, + { + "id": "51392ca4-f824-42d8-8808-8110ec8e0e22", + "name": "Main Street Philharmonic at Main Street, U.S.A.", + "entityType": "SHOW", + "coverage": { + "firstRecordedAt": "2021-07-04" + }, + "days": [ + { + "date": "2026-08-01", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-02", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 2, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-04", + "firstOperatingAt": "2026-08-04T04:05:15Z", + "lastClosedAt": null, + "operatingMinutes": 1110, + "downMinutes": 0, + "unknownMinutes": 324, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-05", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-06", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-07", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-08", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-09", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 4, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-11", + "firstOperatingAt": "2026-08-11T04:07:02Z", + "lastClosedAt": null, + "operatingMinutes": 1170, + "downMinutes": 0, + "unknownMinutes": 262, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-12", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-13", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-14", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-15", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-16", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 6, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-18", + "firstOperatingAt": "2026-08-18T04:05:35Z", + "lastClosedAt": null, + "operatingMinutes": 1170, + "downMinutes": 0, + "unknownMinutes": 264, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-19", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-20", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-21", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-22", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-23", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 4, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-25", + "firstOperatingAt": "2026-08-25T04:01:17Z", + "lastClosedAt": null, + "operatingMinutes": 1110, + "downMinutes": 0, + "unknownMinutes": 328, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-26", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-27", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 630, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-28", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-29", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-08-30", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 4, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + } + ] + }, + { + "id": "55bdcccc-217b-416c-b8b0-4b6a87d16179", + "name": "Cinderella's Royal Table", + "entityType": "RESTAURANT", + "coverage": { + "firstRecordedAt": "2021-07-03" + }, + "days": [ + { + "date": "2026-08-02", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-08-04", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-08-09", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-08-18", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-08-22", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-08-28", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-08-30", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + } + ] + } + ], + "next": "https://api.themeparks.wiki/v1/entity/75ea578a-adc8-4116-a54d-dccb60765ef9/history/daily?from=2026-09-01&to=2026-09-20" +} diff --git a/tests/fixtures/mk_park_daily_page2.json b/tests/fixtures/mk_park_daily_page2.json new file mode 100644 index 0000000..35c9699 --- /dev/null +++ b/tests/fixtures/mk_park_daily_page2.json @@ -0,0 +1,1003 @@ +{ + "id": "75ea578a-adc8-4116-a54d-dccb60765ef9", + "name": "Magic Kingdom Park", + "entityType": "PARK", + "parentId": "e957da41-3552-4cf6-b636-5babc5cbc4e5", + "destinationId": "e957da41-3552-4cf6-b636-5babc5cbc4e5", + "timezone": "America/New_York", + "range": { + "from": "2026-09-01", + "to": "2026-09-20" + }, + "entities": [ + { + "id": "f5aad2d4-a419-4384-bd9a-42f86385c750", + "name": "\"it's a small world\"", + "entityType": "ATTRACTION", + "coverage": { + "firstRecordedAt": "2021-07-03" + }, + "days": [ + { + "date": "2026-09-01", + "firstOperatingAt": "2026-09-01T12:31:10Z", + "lastClosedAt": null, + "operatingMinutes": 874, + "downMinutes": 23, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 10 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 845, + "downMinutes": 23, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 10 + } + }, + "changes": 109 + }, + { + "date": "2026-09-02", + "firstOperatingAt": "2026-09-02T12:31:23Z", + "lastClosedAt": "2026-09-03T02:00:22Z", + "operatingMinutes": 808, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 808, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 187 + }, + { + "date": "2026-09-03", + "firstOperatingAt": "2026-09-03T12:30:30Z", + "lastClosedAt": "2026-09-04T02:00:32Z", + "operatingMinutes": 801, + "downMinutes": 9, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 7, + "p90": 15, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 800, + "downMinutes": 9, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 7, + "p90": 15, + "max": 25 + } + }, + "changes": 187 + }, + { + "date": "2026-09-04", + "firstOperatingAt": "2026-09-04T12:30:47Z", + "lastClosedAt": null, + "operatingMinutes": 898, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 7, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 7, + "p90": 10, + "max": 15 + } + }, + "changes": 159 + }, + { + "date": "2026-09-05", + "firstOperatingAt": "2026-09-05T12:30:59Z", + "lastClosedAt": "2026-09-06T03:01:10Z", + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 3, + "standby": { + "min": 5, + "p50": 20, + "mean": 17, + "p90": 30, + "max": 45 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 20, + "mean": 17, + "p90": 30, + "max": 45 + } + }, + "changes": 200 + }, + { + "date": "2026-09-06", + "firstOperatingAt": "2026-09-06T12:31:02Z", + "lastClosedAt": "2026-09-07T02:00:06Z", + "operatingMinutes": 808, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 15, + "mean": 16, + "p90": 25, + "max": 35 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 808, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 16, + "p90": 25, + "max": 35 + } + }, + "changes": 186 + }, + { + "date": "2026-09-07", + "firstOperatingAt": "2026-09-07T12:30:15Z", + "lastClosedAt": "2026-09-08T02:00:21Z", + "operatingMinutes": 809, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 5 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 809, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 5 + } + }, + "changes": 154 + }, + { + "date": "2026-09-08", + "firstOperatingAt": "2026-09-08T12:43:21Z", + "lastClosedAt": null, + "operatingMinutes": 875, + "downMinutes": 22, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 856, + "downMinutes": 12, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 133 + }, + { + "date": "2026-09-09", + "firstOperatingAt": "2026-09-09T12:30:47Z", + "lastClosedAt": null, + "operatingMinutes": 929, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 5, + "p50": 10, + "mean": 12, + "p90": 25, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 929, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 12, + "p90": 25, + "max": 30 + } + }, + "changes": 186 + }, + { + "date": "2026-09-10", + "firstOperatingAt": "2026-09-10T12:31:03Z", + "lastClosedAt": "2026-09-11T03:01:01Z", + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 3, + "standby": { + "min": 0, + "p50": 5, + "mean": 10, + "p90": 20, + "max": 25 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 10, + "p90": 20, + "max": 25 + } + }, + "changes": 193 + }, + { + "date": "2026-09-11", + "firstOperatingAt": "2026-09-11T11:31:11Z", + "lastClosedAt": null, + "operatingMinutes": 908, + "downMinutes": 50, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 878, + "downMinutes": 50, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 131 + }, + { + "date": "2026-09-12", + "firstOperatingAt": "2026-09-12T11:31:22Z", + "lastClosedAt": "2026-09-13T03:00:22Z", + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 30 + } + }, + "changes": 214 + }, + { + "date": "2026-09-13", + "firstOperatingAt": "2026-09-13T11:31:35Z", + "lastClosedAt": null, + "operatingMinutes": 957, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 129 + }, + { + "date": "2026-09-14", + "firstOperatingAt": "2026-09-14T12:30:44Z", + "lastClosedAt": "2026-09-15T03:00:43Z", + "operatingMinutes": 864, + "downMinutes": 5, + "unknownMinutes": 2, + "standby": { + "min": 5, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 40 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 864, + "downMinutes": 5, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 13, + "p90": 25, + "max": 40 + } + }, + "changes": 176 + }, + { + "date": "2026-09-15", + "firstOperatingAt": "2026-09-15T12:30:50Z", + "lastClosedAt": null, + "operatingMinutes": 897, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 869, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 5, + "p90": 5, + "max": 15 + } + }, + "changes": 126 + }, + { + "date": "2026-09-16", + "firstOperatingAt": "2026-09-16T12:30:22Z", + "lastClosedAt": "2026-09-17T02:00:19Z", + "operatingMinutes": 809, + "downMinutes": 0, + "unknownMinutes": 3, + "standby": { + "min": 5, + "p50": 10, + "mean": 11, + "p90": 20, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 809, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 5, + "p50": 10, + "mean": 11, + "p90": 20, + "max": 30 + } + }, + "changes": 189 + }, + { + "date": "2026-09-17", + "firstOperatingAt": "2026-09-17T12:31:24Z", + "lastClosedAt": "2026-09-18T03:00:32Z", + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 10, + "mean": 12, + "p90": 20, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 868, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 10, + "mean": 12, + "p90": 20, + "max": 30 + } + }, + "changes": 185 + }, + { + "date": "2026-09-18", + "firstOperatingAt": "2026-09-18T11:32:34Z", + "lastClosedAt": null, + "operatingMinutes": 956, + "downMinutes": 0, + "unknownMinutes": 1, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 927, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 6, + "p90": 10, + "max": 15 + } + }, + "changes": 146 + }, + { + "date": "2026-09-19", + "firstOperatingAt": "2026-09-19T11:31:47Z", + "lastClosedAt": "2026-09-20T03:00:56Z", + "operatingMinutes": 927, + "downMinutes": 1, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 15, + "mean": 16, + "p90": 30, + "max": 40 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 927, + "downMinutes": 1, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 15, + "mean": 16, + "p90": 30, + "max": 40 + } + }, + "changes": 218 + }, + { + "date": "2026-09-20", + "firstOperatingAt": "2026-09-20T11:31:52Z", + "lastClosedAt": null, + "operatingMinutes": 956, + "downMinutes": 0, + "unknownMinutes": 2, + "standby": { + "min": 0, + "p50": 5, + "mean": 9, + "p90": 15, + "max": 30 + }, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 928, + "downMinutes": 0, + "unknownMinutes": 0, + "standby": { + "min": 0, + "p50": 5, + "mean": 8, + "p90": 15, + "max": 30 + } + }, + "changes": 142 + } + ] + }, + { + "id": "51392ca4-f824-42d8-8808-8110ec8e0e22", + "name": "Main Street Philharmonic at Main Street, U.S.A.", + "entityType": "SHOW", + "coverage": { + "firstRecordedAt": "2021-07-04" + }, + "days": [ + { + "date": "2026-09-01", + "firstOperatingAt": "2026-09-01T04:00:50Z", + "lastClosedAt": null, + "operatingMinutes": 1110, + "downMinutes": 0, + "unknownMinutes": 329, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-02", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 630, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-03", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 630, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-04", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-05", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-06", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 3, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-08", + "firstOperatingAt": "2026-09-08T04:02:07Z", + "lastClosedAt": null, + "operatingMinutes": 1110, + "downMinutes": 0, + "unknownMinutes": 327, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-09", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-10", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-11", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-12", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-13", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 1, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-15", + "firstOperatingAt": "2026-09-15T04:02:56Z", + "lastClosedAt": null, + "operatingMinutes": 1110, + "downMinutes": 0, + "unknownMinutes": 327, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-16", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 630, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-17", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-18", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 4, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-19", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "showCount": 5, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + }, + { + "date": "2026-09-20", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 2, + "showCount": 0, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 0, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 1 + } + ] + }, + { + "id": "55bdcccc-217b-416c-b8b0-4b6a87d16179", + "name": "Cinderella's Royal Table", + "entityType": "RESTAURANT", + "coverage": { + "firstRecordedAt": "2021-07-03" + }, + "days": [ + { + "date": "2026-09-01", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-09-05", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-09-11", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-09-14", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 570, + "inParkHours": { + "scheduledMinutes": 870, + "operatingMinutes": 870, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-09-16", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 630, + "inParkHours": { + "scheduledMinutes": 810, + "operatingMinutes": 810, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + }, + { + "date": "2026-09-18", + "firstOperatingAt": null, + "lastClosedAt": null, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 510, + "inParkHours": { + "scheduledMinutes": 930, + "operatingMinutes": 930, + "downMinutes": 0, + "unknownMinutes": 0 + }, + "changes": 0 + } + ] + } + ], + "next": null +} diff --git a/tests/unit/test_backfill.py b/tests/unit/test_backfill.py index 03b1a1d..61ead2e 100644 --- a/tests/unit/test_backfill.py +++ b/tests/unit/test_backfill.py @@ -20,19 +20,67 @@ from __future__ import annotations import csv +import inspect import json from datetime import date from pathlib import Path +from typing import Union import pytest -from pydantic import ValidationError - -from themeparks import APIError, BudgetExhaustedError, RateLimitError, backfill -from themeparks._ergonomic.history import EntityRef, HistorySpan -from themeparks._generated.models import HistoryDailyRow, HistoryErrorWindowExceeded +from pydantic import BaseModel, ValidationError + +from themeparks import APIError, BudgetExhaustedError, NetworkError, RateLimitError, backfill +from themeparks._client import PACKAGE_VERSION +from themeparks._ergonomic.history import EntityRef, HistoryApi, HistoryPage, HistorySpan +from themeparks._generated.models import ( + HistoryDailyRow, + HistoryDailyStats, + HistoryErrorWindowExceeded, +) from themeparks._transport import _format_error_message from themeparks.backfill import _Park +FIXTURES = Path(__file__).resolve().parents[1] / "fixtures" + +_CSV_FINGERPRINT = backfill._columns_fingerprint("csv") + + +def _state_file(tmp_path: Path, park: str = "p", fmt: str = "ndjson", **fields: object) -> Path: + """Write a state file this build accepts, overriding the fields a test cares about. + + Built from `_record`'s own vocabulary rather than typed out, so a test cannot + silently drift from the contract the code enforces -- which is how a suite of + nine tests once passed against a body shape the API never sends. + """ + path = backfill.state_path_for(tmp_path, park, fmt) + state: dict[str, object] = { + "sdk": backfill.SDK_NAME, + "sdkVersion": PACKAGE_VERSION, + "stateVersion": backfill.STATE_VERSION, + "format": fmt, + "columns": backfill._columns_fingerprint(fmt), + "start": "2025-01-01", + "end": "2026-09-28", + "lastDay": None, + "resumeFrom": None, + "complete": False, + } + state.update(fields) + path.write_text(json.dumps(state), encoding="utf-8") + return path + + +def _next_url(start: str | None) -> str | None: + """A `next` URL in the shape the API sends, or None on the last page. + + The command reads the day out of this URL rather than doing date arithmetic, + so a stub that hands it a bare day would test a code path that does not exist. + """ + if start is None: + return None + return f"https://api.themeparks.wiki/v1/entity/park-1/history/daily?from={start}&to=2026-09-28" + + # CAPTURED FROM PRODUCTION, 2026-09-28: anonymous GET of # /v1/entity/7340550b-c14d-4def-80bb-acdb51d49a66/history/daily starting 2021-07-03. PRODUCTION_403_BODY = { @@ -98,6 +146,12 @@ def __init__(self, archive_from: str, through: str, floor: str | None) -> None: self.floor = floor self.calls: list[str] = [] self.ends: list[str] = [] + # One page ending at `through`, with nothing after it, unless a test says + # otherwise. Each entry is (this page's last day, the next page's start). + self.entity_name = "Test Coaster" + self.pages: list[tuple[str, str | None]] = [(through, None)] + #: Day of the page after which the hourly budget runs out, or None. + self.budget_after_page: str | None = None def span(self) -> HistorySpan: return HistorySpan( @@ -106,19 +160,32 @@ def span(self) -> HistorySpan: date.fromisoformat(self.through), ) - def days_with_entities(self, start, end): + def days_with_entities(self, start=None, end=None, *, max_wait=120.0, on_page=None): """Mirrors the real signature: the NAME comes from the response. Not `days()`. The command switched to `days_with_entities` so a row is labelled with the name the history response gave for it, rather than the park's current children list -- rides get renamed and old rows must keep the name they were recorded under. + + `pages` drives the paging: one entry per page, `(last day it covered, the + day the next page starts on or None)`. The default is one page, so a test + that does not care about paging does not have to say so. `on_page` fires + AFTER that page's rows, which is the contract the checkpoint depends on. """ self.calls.append(str(start)) self.ends.append(str(end)) if self.floor is not None and str(start) < self.floor: raise _window_403(self.floor) - yield (EntityRef("ent-1", "Test Coaster", "ATTRACTION"), _row(str(start))) + for covered, next_from in self.pages: + yield (EntityRef("ent-1", self.entity_name, "ATTRACTION"), _row(covered)) + if on_page is not None: + on_page(HistoryPage(str(start), covered, _next_url(next_from))) + # The budget is spent at a PAGE BOUNDARY, which is the only place a + # real one can be: rows come from a fully parsed envelope, so the + # next request is what gets refused. + if self.budget_after_page == covered: + raise BudgetExhaustedError("429", status=429, body={}, url="u", retry_after=2700.0) class _Entity: @@ -237,7 +304,7 @@ def test_does_not_retry_once_rows_are_written(self, tmp_path: Path) -> None: hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) calls = {"n": 0} - def days_with_entities(start, end): + def days_with_entities(start=None, end=None, *, max_wait=120.0, on_page=None): calls["n"] += 1 yield (EntityRef("ent-1", "Test Coaster", "ATTRACTION"), _row("2025-01-01")) raise _window_403("2025-08-25") @@ -287,14 +354,14 @@ def test_exactly_one_header_even_through_the_403_recovery(self, tmp_path: Path) # row as data and every numeric column came back as text. hist = _History(archive_from="2021-07-03", through="2026-09-28", floor="2025-08-25") backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "csv") - text = (tmp_path / "p.csv").read_text(encoding="utf-8") + text = (tmp_path / "p.csv").read_text(encoding="utf-8-sig") assert sum(1 for line in text.splitlines() if line.startswith("parkId,")) == 1 def test_identity_columns_come_first_and_carry_the_response_name(self, tmp_path: Path) -> None: hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) backfill.backfill_park(_Client(hist), _Park("park-id", "Park Name"), tmp_path, "csv") rows = list( - csv.DictReader((tmp_path / "park-id.csv").read_text(encoding="utf-8").splitlines()) + csv.DictReader((tmp_path / "park-id.csv").read_text(encoding="utf-8-sig").splitlines()) ) assert rows[0]["parkId"] == "park-id" assert rows[0]["parkName"] == "Park Name" @@ -336,7 +403,8 @@ def test_completion_is_recorded_not_inferred(self, tmp_path: Path) -> None: # as "never started". That ambiguity is what let a retry re-download a # finished park in full. self._run(tmp_path) - state = json.loads((tmp_path / f"p{backfill.STATE_SUFFIX}").read_text(encoding="utf-8")) + state_file = backfill.state_path_for(tmp_path, "p", "ndjson") + state = json.loads(state_file.read_text(encoding="utf-8")) assert state["complete"] is True assert state["format"] == "ndjson" assert state["start"] == "2025-01-01" @@ -371,7 +439,9 @@ def test_every_key_is_asserted_not_just_non_emptiness(self, tmp_path: Path) -> N hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) backfill.backfill_park(_Client(hist), _Park("park-id", "Park Name"), tmp_path, "ndjson") line = json.loads((tmp_path / "park-id.ndjson").read_text(encoding="utf-8").splitlines()[0]) - row = _row("2025-01-01") + # The stub dates its row at the last day of the page it belongs to, which + # is what a real response does: a page's rows run to `range.to`. + row = _row("2026-09-28") assert line == { "parkId": "park-id", "parkName": "Park Name", @@ -388,7 +458,7 @@ class TestBudgetExhaustion: def _hist_that_runs_out(self) -> _History: hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) - def days_with_entities(start, end): + def days_with_entities(start=None, end=None, *, max_wait=120.0, on_page=None): hist.calls.append(str(start)) # DELIBERATELY out of order, and the second entity's days end EARLIER. # That is the real shape: `_daily_rows` walks entities and then each @@ -412,12 +482,18 @@ def test_returns_ex_tempfail_so_a_scheduler_retries(self, tmp_path: Path) -> Non def test_records_the_furthest_day_so_a_rerun_continues(self, tmp_path: Path) -> None: hist = self._hist_that_runs_out() backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") - state = json.loads((tmp_path / f"p{backfill.STATE_SUFFIX}").read_text(encoding="utf-8")) + state_file = backfill.state_path_for(tmp_path, "p", "ndjson") + state = json.loads(state_file.read_text(encoding="utf-8")) assert state["complete"] is False # The MAX day seen, not the last row yielded. The stub's final row is # 2025-02-14, so recording "last" instead of "max" would rewind the resume # point by three weeks and re-download them. - assert state["last_day"] == "2025-03-09" + assert state["lastDay"] == "2025-03-09" + # The budget died part-way through the FIRST page, so no page boundary was + # ever reported and there is nothing exact to resume from. `last_day` is + # the documented fallback for precisely this: one duplicated day, rather + # than starting from the top and appending a second copy of everything. + assert state["resumeFrom"] is None def test_a_spent_budget_on_the_coverage_call_also_returns_75(self, tmp_path: Path) -> None: # span() is the FIRST request a resumed run makes, while the hourly window @@ -442,3 +518,747 @@ def test_skips_without_making_a_request_at_all(self, tmp_path: Path, capsys) -> assert code == 0 assert hist.calls == [], "asked the API for a range it had already ruled out" assert "nothing in your window" in capsys.readouterr().err + + +class TestStubsMatchTheRealSdk: + """A stand-in that has drifted from the real object proves nothing. + + This is the fifth time a hand-written double and the code it stood in for + disagreed, and one of those shipped: a 403 handler written from a traceback + read a body shape the API never sends, and nine tests passed against a + fixture retyped from the same traceback. A stub whose signature is checked + against the real method cannot silently accept a call the SDK would reject. + """ + + def test_the_history_stub_takes_what_the_real_method_takes(self) -> None: + real = inspect.signature(HistoryApi.days_with_entities) + stub = inspect.signature(_History.days_with_entities) + real_params = [p for name, p in real.parameters.items() if name != "self"] + stub_params = [p for name, p in stub.parameters.items() if name != "self"] + assert [p.name for p in stub_params] == [p.name for p in real_params] + assert [p.kind for p in stub_params] == [p.kind for p in real_params] + + def test_the_stub_refuses_a_call_the_real_method_would_refuse(self) -> None: + # Guard the guard: if the stub swallowed **kwargs the check above passes + # while the stub accepts anything, which is the failure it exists to stop. + hist = _History(archive_from="2025-01-01", through="2026-01-01", floor=None) + with pytest.raises(TypeError): + list(hist.days_with_entities("2025-01-01", "2026-01-01", nonsense=True)) + + +class TestResumeCheckpoint: + """Where a rerun carries on from. It was the newest ROW, which duplicates.""" + + def _paged(self) -> _History: + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + # Two pages, the real shape: the first covers through 2025-01-31 and the + # server says carry on at 2025-02-01. The budget then runs out on the + # second page, after the first has been written in full. + hist.pages = [("2025-01-31", "2025-02-01"), ("2025-03-02", None)] + hist.raise_on_page = None + return hist + + def test_records_the_day_the_next_page_starts_on(self, tmp_path: Path) -> None: + # THE HEADLINE FIX OF THIS BRANCH, AND IT HAD NO TEST. The version of this + # test that shipped asserted only `complete is True` -- which another test + # already covered -- and could not observe the boundary at all, because the + # completion path hard-wires resumeFrom=None. Mutating the checkpoint to + # record the newest row instead of the page boundary left all 301 tests + # green. + # + # Page one covers through 2025-01-31 and the server says carry on at + # 2025-02-01; the one entity's newest row is 2025-01-31. The budget is then + # spent, so the state file is what a rerun reads, and the three values must + # stay distinguishable. + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + hist.pages = [("2025-01-31", "2025-02-01"), ("2025-03-02", None)] + hist.budget_after_page = "2025-01-31" + + code = backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + + assert code == backfill.EX_TEMPFAIL + state = json.loads( + backfill.state_path_for(tmp_path, "p", "ndjson").read_text(encoding="utf-8") + ) + assert state["resumeFrom"] == "2025-02-01", "checkpointed the row, not the boundary" + assert state["lastDay"] == "2025-01-31" + assert state["complete"] is False + # One row written, from page one only: page two was never served. + assert len((tmp_path / "p.ndjson").read_text(encoding="utf-8").strip().splitlines()) == 1 + + def test_a_rerun_asks_for_the_page_boundary_not_the_newest_row(self, tmp_path: Path) -> None: + # THE DEFECT. An entity that stopped reporting has no rows for the tail + # days of its page, so the newest row is earlier than the page covered. + # Resuming there re-fetches days already in the file and appends every row + # of them again, breaking the (entityId, date) key the file is documented + # to have -- on the exit-75 path, which is the ordinary path for a long + # back fill rather than an edge case. + (tmp_path / "p.ndjson").write_text('{"a": 1}\n', encoding="utf-8") + _state_file(tmp_path, lastDay="2025-01-28", resumeFrom="2025-02-01") + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + assert hist.calls == ["2025-02-01"], "resumed from the newest row, not the boundary" + + def test_falls_back_to_last_day_for_a_state_file_without_a_boundary( + self, tmp_path: Path + ) -> None: + # 3.3.0 wrote no resume_from, and a run that dies inside its first page + # never reports one. One duplicated day beats starting from the top and + # appending a second copy of the whole archive. + (tmp_path / "p.ndjson").write_text('{"a": 1}\n', encoding="utf-8") + # No boundary recorded: a run that died inside its first page. + _state_file(tmp_path, lastDay="2025-01-28", resumeFrom=None) + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + assert hist.calls == ["2025-01-28"] + + def test_the_boundary_is_read_from_the_servers_own_url(self) -> None: + # No date arithmetic anywhere: the server says where the next page starts + # and that string is what gets recorded. + url = ( + "https://api.themeparks.wiki/v1/entity/75ea578a-adc8-4116-a54d-dccb60765ef9" + "/history/daily?from=2026-09-01&to=2026-09-20" + ) + assert backfill._next_page_start(url) == "2026-09-01" + assert backfill._next_page_start(None) is None + assert backfill._next_page_start("") is None + assert backfill._next_page_start("https://api.themeparks.wiki/v1/x") is None + + +class TestEveryFieldReachesTheFile: + """The CSV carried 26 of the 41 columns the schema defines.""" + + def test_columns_are_derived_from_the_model(self) -> None: + # Hand-typed, the list drifted three ways at once: unknownMinutes and the + # whole inParkHours block missing, extremeWaits missing, and singleRider + # carrying two of its five percentiles while standby carried all five. + # Ten of thirty-six data fields absent from a file people pay for. + for name in ("unknownMinutes", "extremeWaitsStandby", "inParkHoursScheduledMinutes"): + assert name in backfill.CSV_COLUMNS, name + for prefix in ("standby", "singleRider", "inParkHoursStandby", "inParkHoursSingleRider"): + stats = [c[len(prefix) :] for c in backfill.DATA_COLUMNS if c.startswith(prefix)] + assert stats[:5] == ["Min", "P50", "Mean", "P90", "Max"], prefix + + def test_every_field_in_a_real_page_of_rows_has_a_column(self) -> None: + """THE DRIFT GATE, with an oracle that is not the code. + + The version this replaces computed both sides of its assertion with the + production helper `_stats_model` and compared them: two runs of one + algorithm over one input. Making `_stats_model` return None + unconditionally, which drops 30 of 36 columns, left it GREEN. Its + docstring called it the drift gate. + + The oracle here is 108 rows of a real capture, flattened by an independent + walk written in terms of the JSON rather than the model. + """ + rows = [ + day + for page in ("mk_park_daily_page1.json", "mk_park_daily_page2.json") + for entity in json.loads((FIXTURES / page).read_text(encoding="utf-8"))["entities"] + for day in entity["days"] + ] + assert len(rows) == 108, "the capture changed; recount before trusting this" + + def flat(obj: dict[str, object], prefix: str = "") -> list[str]: + out: list[str] = [] + for key, value in obj.items(): + name = key if not prefix else f"{prefix}{key[0].upper()}{key[1:]}" + if isinstance(value, dict): + out.extend(flat(value, name)) + else: + out.append(name) + return out + + seen = {name for row in rows for name in flat(row)} + assert seen, "flattened nothing" + missing = sorted(seen - set(backfill.CSV_COLUMNS)) + assert missing == [], f"the API sends these and the CSV has no column: {missing}" + # The generated list is a superset, not a coincidence: this park has no + # single-rider queue, so those columns are in the header and in no row. + assert "singleRiderP90" in backfill.CSV_COLUMNS + assert "singleRiderP90" not in seen + assert len(backfill.DATA_COLUMNS) == 36 + + def test_an_absent_stats_block_is_empty_cells_never_zeroes(self, tmp_path: Path) -> None: + # A zero would read as "measured, and it was nothing". Absent means no + # wait was in force, or the park published no hours that day. + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "csv") + rows = list(csv.DictReader((tmp_path / "p.csv").open(encoding="utf-8-sig"))) + assert rows[0]["inParkHoursStandbyP50"] == "" + assert rows[0]["extremeWaitsStandby"] == "" + assert rows[0]["entityName"] == "Test Coaster" + + def test_the_header_is_the_column_list(self, tmp_path: Path) -> None: + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "csv") + header = (tmp_path / "p.csv").read_text(encoding="utf-8-sig").splitlines()[0] + assert header.split(",") == backfill.CSV_COLUMNS + + +class TestTheCommandIdentifiesItself: + def test_the_user_agent_names_the_command_and_the_sdk(self) -> None: + # It was the literal "themeparks-backfill/1": a hardcoded 1 that could + # never match a release, and it REPLACED the SDK's user agent, so a + # support question about a bad download had no version at either end. + assert ( + f"themeparks-backfill/{PACKAGE_VERSION} themeparks-sdk-py/{PACKAGE_VERSION}" + ) == backfill.USER_AGENT + assert "/1" not in backfill.USER_AGENT.replace(PACKAGE_VERSION, "") + + def test_version_prints_the_package_version(self, capsys) -> None: + with pytest.raises(SystemExit) as caught: + backfill.main(["--version"]) + assert caught.value.code == 0 + assert PACKAGE_VERSION in capsys.readouterr().out + + +class TestFailureIsNotATraceback: + """A customer who has just paid reads a traceback as the tool being broken.""" + + def test_an_unreachable_api_is_resumable_so_exit_75(self, capsys, monkeypatch) -> None: + # A connection reset IS resumable, so a scheduler should retry rather than + # alert. Python exited 1 here while the JavaScript SDK exited 75: opposite + # semantics for one event, on the number a cron acts on. + def boom(argv=None): + raise NetworkError("connection refused") + + monkeypatch.setattr(backfill, "main", boom) + assert backfill.cli() == backfill.EX_TEMPFAIL + err = capsys.readouterr().err + assert "connection refused" in err + assert "resumable" in err + assert "Traceback" not in err + + def test_something_the_api_rejected_is_exit_1(self, capsys, monkeypatch) -> None: + # The other half, so the rule is not "every failure is retryable". A bad + # range or a revoked key will fail identically next time. + def boom(argv=None): + raise APIError("400 Bad Request", status=400, body={}, url="u") + + monkeypatch.setattr(backfill, "main", boom) + assert backfill.cli() == 1 + err = capsys.readouterr().err + assert "400" in err + assert "resumable" not in err + + def test_ctrl_c_says_how_to_continue(self, capsys, monkeypatch) -> None: + def stop(argv=None): + raise KeyboardInterrupt + + monkeypatch.setattr(backfill, "main", stop) + assert backfill.cli() == 130 + assert "run the same command again" in capsys.readouterr().err.lower() + + def test_a_failed_park_leaves_no_empty_file(self, tmp_path: Path) -> None: + # Opening the file created it before the first request, so a park that + # failed with nothing written left a 0-byte file that reads as "this park + # has no history". + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + + def days_with_entities(start=None, end=None, *, max_wait=120.0, on_page=None): + raise APIError("500 Server Error", status=500, body={}, url="u") + yield # pragma: no cover - never reached, keeps this a generator + + hist.days_with_entities = days_with_entities # type: ignore[assignment] + with pytest.raises(APIError): + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + assert not (tmp_path / "p.ndjson").exists(), "left a 0-byte file behind" + + +class TestListing: + def test_the_destination_total_is_counted_before_the_filter(self, capsys) -> None: + # `--list epcot` printed "all 1 parks" for a destination with six, on the + # one line whose entire job is that number -- and that line tells the + # reader to pass the destination id, so the number is what they act on. + catalogue = [ + ("p1", "EPCOT", "d1", "Walt Disney World Resort"), + ("p2", "Magic Kingdom Park", "d1", "Walt Disney World Resort"), + ("p3", "Disney's Hollywood Studios", "d1", "Walt Disney World Resort"), + ] + assert backfill._print_list(catalogue, "EPCOT") == 0 + out = capsys.readouterr().out + assert "all 3 parks (1 shown)" in out + + def test_nothing_matching_is_exit_1(self, capsys) -> None: + assert backfill._print_list([("p1", "EPCOT", "d1", "WDW")], "zzz") == 1 + + +class TestAmbiguousNames: + def test_an_exact_name_two_parks_share_lists_only_those_two(self) -> None: + # Widening to substrings adds Hong Kong Disneyland Park, which is not what + # was typed, to the one list whose job is "which of these did you mean". + catalogue = [ + ("p1", "Disneyland Park", "d1", "Disneyland Resort"), + ("p2", "Disneyland Park", "d2", "Disneyland Paris"), + ("p3", "Hong Kong Disneyland Park", "d3", "Hong Kong Disneyland Parks"), + ] + with pytest.raises(SystemExit) as caught: + backfill._resolve(catalogue, "Disneyland Park") + message = str(caught.value) + assert "matches 2" in message + assert "Hong Kong" not in message + assert "Disneyland Resort" in message and "Disneyland Paris" in message + + +class TestOneParkFailingIsNotTheRunFailing: + """A 500 on park three used to abandon parks four, five and six.""" + + class _Args: + def __init__(self, out: Path) -> None: + self.out = out + self.format = "ndjson" + self.overwrite = False + + def test_every_park_is_tried_and_the_failures_are_named( + self, tmp_path: Path, capsys, monkeypatch + ) -> None: + attempted: list[str] = [] + + def fake_backfill(tp, park, out_dir, fmt, overwrite=False): + attempted.append(park.id) + if park.id == "p2": + raise APIError("500 Server Error", status=500, body={}, url="u") + return 0 + + monkeypatch.setattr(backfill, "backfill_park", fake_backfill) + targets = [("p1", "One"), ("p2", "Two"), ("p3", "Three")] + code = backfill._run_all(None, targets, self._Args(tmp_path)) + assert attempted == ["p1", "p2", "p3"], "stopped at the first failure" + assert code == 1 + err = capsys.readouterr().err + assert "1 of 3 did not finish: Two" in err + + def test_a_spent_budget_stops_the_whole_run(self, tmp_path: Path, monkeypatch) -> None: + # The opposite rule, and it matters: the next park would spend the + # retry-after on a 429 for nothing, and every state file already says + # where it got to. + attempted: list[str] = [] + + def fake_backfill(tp, park, out_dir, fmt, overwrite=False): + attempted.append(park.id) + return backfill.EX_TEMPFAIL + + monkeypatch.setattr(backfill, "backfill_park", fake_backfill) + code = backfill._run_all(None, [("p1", "One"), ("p2", "Two")], self._Args(tmp_path)) + assert code == backfill.EX_TEMPFAIL + assert attempted == ["p1"] + + def test_all_good_is_exit_0_and_says_nothing(self, tmp_path: Path, capsys, monkeypatch) -> None: + monkeypatch.setattr(backfill, "backfill_park", lambda *a, **k: 0) + code = backfill._run_all(None, [("p1", "One"), ("p2", "Two")], self._Args(tmp_path)) + assert code == 0 + assert "did not finish" not in capsys.readouterr().err + + +class TestTimestampsMatchTheApi: + def test_utc_is_written_the_way_the_api_sends_it(self, tmp_path: Path) -> None: + # `+00:00` and `Z` are the same instant and not the same string. The API + # sends `Z`, the JavaScript SDK's identical command passes it through, and + # Python's isoformat() turned it into `+00:00` -- so the same park exported + # in two languages came back as two different files, 39,201 lines apart on + # a five-year EPCOT run, for no difference in meaning. + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "csv") + row = next(csv.DictReader((tmp_path / "p.csv").open(encoding="utf-8-sig"))) + assert row["firstOperatingAt"].endswith("Z") + assert "+00:00" not in row["firstOperatingAt"] + + def test_a_plain_date_stays_a_plain_date(self) -> None: + assert backfill._scalar(date(2026, 9, 28)) == "2026-09-28" + assert backfill._scalar(None) == "" + + +class TestAnEarlierRunsRowsAreNeverDeleted: + """`written == 0` means "this process wrote nothing", not "the file is empty". + + Both deletions below were added the same day to close an "empty file is a lie" + finding, and both destroyed a resumed customer's archive while closing it. The + state file survived pointing into the middle of the range, so the next run + appended only the tail and recorded `complete: true`. + """ + + def _partial(self, tmp_path: Path) -> None: + (tmp_path / "p.ndjson").write_text('{"row": 1}\n{"row": 2}\n', encoding="utf-8") + _state_file(tmp_path, lastDay="2025-06-30", resumeFrom="2025-07-01") + + def test_a_transient_failure_on_a_resumed_run_keeps_the_file(self, tmp_path: Path) -> None: + self._partial(tmp_path) + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + + def days_with_entities(start=None, end=None, *, max_wait=120.0, on_page=None): + raise NetworkError("connection reset by peer") + yield # pragma: no cover - keeps this a generator + + hist.days_with_entities = days_with_entities # type: ignore[assignment] + with pytest.raises(NetworkError): + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + + assert (tmp_path / "p.ndjson").exists(), "destroyed rows an earlier run downloaded" + assert (tmp_path / "p.ndjson").read_text(encoding="utf-8").count("\n") == 2 + + def test_a_closed_window_on_a_resumed_run_keeps_the_file_and_fails( + self, tmp_path: Path, capsys + ) -> None: + # A key rotated out of a scheduler's environment, or a lapsed subscription: + # the plan no longer reaches the resume point. Deleting the file and + # exiting 0 made the scheduler log success and left a permanent trap -- the + # data gone, the state surviving, every later run re-entering the branch. + self._partial(tmp_path) + hist = _History(archive_from="2025-01-01", through="2025-02-01", floor=None) + code = backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + + assert code == 1, "exit 0 tells a scheduler this succeeded" + assert (tmp_path / "p.ndjson").exists(), "destroyed rows an earlier run downloaded" + assert "left alone" in capsys.readouterr().err + + def test_a_first_run_that_fails_with_nothing_written_leaves_no_file( + self, tmp_path: Path + ) -> None: + # The other half of the rule, so the guard cannot be "never delete". + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + + def days_with_entities(start=None, end=None, *, max_wait=120.0, on_page=None): + raise NetworkError("connection reset by peer") + yield # pragma: no cover + + hist.days_with_entities = days_with_entities # type: ignore[assignment] + with pytest.raises(NetworkError): + backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + assert not (tmp_path / "p.ndjson").exists() + + +class TestAStateFileThisBuildCannotResume: + """Refuse, never resume on a guess. Each of these corrupted a file instead.""" + + def _rows(self, tmp_path: Path, name: str = "p.csv") -> None: + (tmp_path / name).write_text("old,header\n1,2\n", encoding="utf-8") + + def test_a_different_column_layout_is_refused(self, tmp_path: Path, capsys) -> None: + # 3.3.0 wrote 19 columns in a different order; this build writes 41. The + # old code compared only the format, so 41-field rows were appended under + # a 19-column header: pandas refuses the file, DictReader reads standbyMin + # as changes, exit 0 either way. The fingerprint is what makes it visible. + self._rows(tmp_path) + _state_file(tmp_path, fmt="csv", columns="0000deadbeef0000", lastDay="2025-06-30") + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + code = backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "csv") + assert code == 1 + err = capsys.readouterr().err + assert "column layout changed" in err + assert "--overwrite" in err + assert (tmp_path / "p.csv").read_text(encoding="utf-8-sig") == "old,header\n1,2\n" + + def test_a_state_file_from_the_other_sdk_is_refused(self, tmp_path: Path, capsys) -> None: + # Same filename, and `format`/`start`/`end`/`complete` spelled identically, + # so the safe paths interoperated and nothing warned -- while a Python run + # interrupted at 64 rows and resumed by the JavaScript command produced 172 + # rows, 64 of them duplicates, marked complete. + self._rows(tmp_path, "p.ndjson") + _state_file(tmp_path, sdk="js", lastDay="2025-06-30", resumeFrom="2025-07-01") + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + code = backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") + assert code == 1 + assert "js SDK" in capsys.readouterr().err + + def test_a_future_state_version_is_refused(self, tmp_path: Path, capsys) -> None: + self._rows(tmp_path, "p.ndjson") + _state_file(tmp_path, stateVersion=backfill.STATE_VERSION + 1, lastDay="2025-06-30") + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + assert backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "ndjson") == 1 + assert "different version" in capsys.readouterr().err + + def test_the_fingerprint_tracks_the_real_header(self) -> None: + # Guard the guard: a constant fingerprint would refuse nothing. + assert backfill._columns_fingerprint("csv") != backfill._columns_fingerprint("ndjson") + assert len(backfill._columns_fingerprint("csv")) == 16 + original = backfill.CSV_COLUMNS + try: + backfill.CSV_COLUMNS = [*original, "aNewColumn"] # type: ignore[misc] + assert backfill._columns_fingerprint("csv") != _CSV_FINGERPRINT + finally: + backfill.CSV_COLUMNS = original # type: ignore[misc] + + +class TestFormatsDoNotShareOneStateFile: + def test_ndjson_csv_ndjson_does_not_double_the_first_file(self, tmp_path: Path) -> None: + # One state file served both formats, so a round trip left the ndjson file + # with no state describing it: resuming false, has_rows false, opened in + # append mode. Every (entityId, date) duplicated, exit 0, "done". Verbatim + # the defect the state file was introduced to prevent. + def run(fmt: str) -> int: + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + return backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, fmt) + + assert run("ndjson") == 0 + first = (tmp_path / "p.ndjson").read_text(encoding="utf-8") + assert run("csv") == 0 + assert run("ndjson") == 0 + again = (tmp_path / "p.ndjson").read_text(encoding="utf-8") + assert again == first, "appended a second copy" + + def test_each_format_gets_its_own_state_file(self, tmp_path: Path) -> None: + assert backfill.state_path_for(tmp_path, "p", "csv") != backfill.state_path_for( + tmp_path, "p", "ndjson" + ) + assert backfill.state_path_for(tmp_path, "p", "csv").name.startswith("p.csv") + + +class TestTheSpreadsheetIsTheReader: + """The CSV's primary consumer is Excel, and it is a paid deliverable.""" + + def _write(self, tmp_path: Path, park_name: str = "P", entity: str = "Test Coaster") -> Path: + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + hist.entity_name = entity + backfill.backfill_park(_Client(hist), _Park("p", park_name), tmp_path, "csv") + return tmp_path / "p.csv" + + def test_the_file_starts_with_a_utf8_bom(self, tmp_path: Path) -> None: + # Without it Excel on Windows reads the local code page and renders + # `Walt Disney World® Resort` as mojibake. + raw = self._write(tmp_path, park_name="Walt Disney World® Resort").read_bytes() + assert raw.startswith(b"\xef\xbb\xbf"), "no BOM: Excel will mangle every ® and accent" + assert raw.count(b"\xef\xbb\xbf") == 1, "a BOM per row, or per resume, is not a BOM" + assert "Walt Disney World® Resort" in raw.decode("utf-8-sig") + + def test_a_resumed_file_does_not_gain_a_second_bom_or_header(self, tmp_path: Path) -> None: + self._write(tmp_path) + _state_file(tmp_path, fmt="csv", lastDay="2025-06-30", resumeFrom="2025-07-01") + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + assert backfill.backfill_park(_Client(hist), _Park("p", "P"), tmp_path, "csv") == 0 + raw = (tmp_path / "p.csv").read_bytes() + assert raw.count(b"\xef\xbb\xbf") == 1 + assert raw.decode("utf-8-sig").count("parkId,") == 1, "a resumed CSV gained a second header" + + def test_a_name_a_spreadsheet_would_execute_is_defused(self, tmp_path: Path) -> None: + rows = list( + csv.DictReader( + self._write(tmp_path, entity="=cmd|' /C calc'!A0").open(encoding="utf-8-sig") + ) + ) + assert rows[0]["entityName"].startswith("'="), "Excel would run this as a formula" + + def test_a_negative_number_stays_a_number(self, tmp_path: Path) -> None: + # The reason the guard is not a bare startswith: prefixing `-5` would turn + # every negative value in the file into text. + assert backfill._defuse("-5") == "-5" + assert backfill._defuse("-5.25") == "-5.25" + assert backfill._defuse("+1") == "+1" + assert backfill._defuse("@SUM(1)") == "'@SUM(1)" + + def test_a_carriage_return_in_a_name_does_not_split_the_row(self, tmp_path: Path) -> None: + # A bare CR unquoted makes one row parse as two, with every later column + # shifted. The JS regex omitted \r; Python's csv module quotes it. + path = self._write(tmp_path, entity="Space Mountain\rFastPass") + # READ WITH A REAL CSV READER over the file. `splitlines()` splits on a bare + # CR whether or not it is inside quotes, so reading that way tests Python's + # string method rather than the file. + with path.open(encoding="utf-8-sig", newline="") as handle: + rows = list(csv.reader(handle)) + assert len(rows) == 2, "the row split" + assert len(rows[1]) == len(backfill.CSV_COLUMNS) + assert rows[1][3] == "Space Mountain\rFastPass", "the name was altered" + # And the BYTES: the field is quoted, so no reader can split it. Read as + # bytes, because text mode translates the CR to a newline before any + # assertion can see it. + assert b'"Space Mountain\rFastPass"' in path.read_bytes() + + +class TestTheRowDoesNotRenameTheRun: + def test_an_undeclared_identity_field_on_a_row_does_not_win(self, tmp_path: Path) -> None: + # Models keep undeclared fields now, so a row carrying its own `entityId` + # or `parkName` would win the merge and silently relabel every line of a + # paid export. Before `extra="allow"` such a field was dropped at parse + # time, so identity always survived; the fix and the hazard arrived + # together. + hist = _History(archive_from="2025-01-01", through="2026-09-28", floor=None) + hijacked = HistoryDailyRow.model_validate( + { + "date": "2026-09-28", + "firstOperatingAt": None, + "lastClosedAt": None, + "operatingMinutes": 1, + "downMinutes": 0, + "changes": 1, + "entityId": "FROM-THE-ROW", + "parkName": "FROM-THE-ROW", + } + ) + assert hijacked.model_extra == {"entityId": "FROM-THE-ROW", "parkName": "FROM-THE-ROW"} + + def days_with_entities(start=None, end=None, *, max_wait=120.0, on_page=None): + yield (EntityRef("real-guid", "Real Ride", "ATTRACTION"), hijacked) + + hist.days_with_entities = days_with_entities # type: ignore[assignment] + backfill.backfill_park(_Client(hist), _Park("real-park", "Real Park"), tmp_path, "ndjson") + line = json.loads( + (tmp_path / "real-park.ndjson").read_text(encoding="utf-8").splitlines()[0] + ) + assert line["entityId"] == "real-guid" + assert line["parkName"] == "Real Park" + # Identity still reads first: a dict keeps the position of the first insert. + assert list(line)[:5] == backfill.IDENTITY_COLUMNS + + +class TestTheColumnDerivationCannotBeFooled: + def test_a_list_of_models_is_not_a_nested_block(self) -> None: + # `get_args(list[Stats])` is `(Stats,)`, so the model was found and the + # annotation flattened as one object: phantom columns in the header, then + # AttributeError on the first row. Any array-of-objects the API adds. + assert backfill._stats_model(Union[HistoryDailyStats, None]) is HistoryDailyStats + for collection in ( + list[HistoryDailyStats], + dict[str, HistoryDailyStats], + set[str], + tuple[HistoryDailyStats, ...], + ): + assert backfill._stats_model(collection) is None, collection + + def test_duplicate_column_names_are_impossible(self) -> None: + # The rule is not injective. A collision writes one value + # into two slots under a right-looking header, so it is an import-time + # failure rather than a wrong number in a customer's file. + assert len(set(backfill.CSV_COLUMNS)) == len(backfill.CSV_COLUMNS) + + +def _real_catalogue() -> list[tuple[str, str, str, str]]: + """Every park in the real 101-destination capture, as `_catalogue` builds it. + + The resolution tests used three- and four-row catalogues invented in this file, + so every mutation of the matching logic survived -- while `destinations.json`, + which holds every character case that has ever broken it, sat in the fixtures + directory referenced only by its own README. + """ + return [ + (park["id"], park["name"], dest["id"], dest["name"]) + for dest in json.loads((FIXTURES / "destinations.json").read_text(encoding="utf-8"))[ + "destinations" + ] + for park in (dest.get("parks") or []) + ] + + +class TestResolutionAgainstTheRealCatalogue: + CATALOGUE = _real_catalogue() + WDW = "e957da41-3552-4cf6-b636-5babc5cbc4e5" + MK = "75ea578a-adc8-4116-a54d-dccb60765ef9" + + def test_the_capture_is_what_the_readme_says(self) -> None: + # Guard the oracle: if this shrinks, every test below weakens silently. + assert len(self.CATALOGUE) == 127 + names = {c[3] for c in self.CATALOGUE} + assert "Walt Disney World® Resort" in names + assert "Walibi Rhône-Alpes" in names + assert sum(1 for c in self.CATALOGUE if c[1] == "Disneyland Park") == 2 + + def test_a_name_carrying_a_registered_mark_resolves_without_it(self) -> None: + # The documented example did not work: matching used bare casefold() and + # the live name is `Walt Disney World® Resort`. + assert len(backfill._resolve(self.CATALOGUE, "Walt Disney World Resort")) == 6 + assert len(backfill._resolve(self.CATALOGUE, "walt-disney-world-resort")) == 6 + assert len(backfill._resolve(self.CATALOGUE, "Walibi Rhone Alpes")) == 1 + + def test_a_unique_substring_resolves_to_the_bare_park_name(self) -> None: + # THE 3.3.0 CORRUPTION. This returned the formatted display label, so the + # parkName column of ~73,000 rows read + # "Magic Kingdom Park (Walt Disney World® Resort)". Four live names reach + # this path. + for query, expected in ( + ("magic kingdom", "Magic Kingdom Park"), + ("PortAventura", "PortAventura Park"), + ("Toverland", "Attractiepark Toverland"), + ("Parque Warner", "Parque Warner Madrid"), + ): + resolved = backfill._resolve(self.CATALOGUE, query) + assert len(resolved) == 1, query + assert resolved[0][1] == expected, f"{query} -> {resolved[0][1]!r}" + assert "(" not in resolved[0][1] + + def test_an_exact_destination_name_beats_an_exact_park_name(self) -> None: + # Seven real names are BOTH a destination and one of that destination's + # several parks -- "Cedar Point" is the destination and its headline park, + # which also has a water park. The destination branch must win, or + # `themeparks-backfill "Cedar Point"` quietly downloads one of the two and + # looks like it worked. The names that share a single-park destination + # cannot tell the two branches apart, which is why this picks a + # multi-park one. + for name, expected in ( + ("Cedar Point", 2), + ("Knott's Berry Farm", 2), + ("Europa-Park", 3), + ("Six Flags Over Texas", 2), + ): + resolved = backfill._resolve(self.CATALOGUE, name) + assert len(resolved) == expected, f"{name} -> {[r[1] for r in resolved]}" + + def test_a_destination_id_expands_and_a_park_id_does_not(self) -> None: + assert len(backfill._resolve(self.CATALOGUE, self.WDW)) == 6 + assert backfill._resolve(self.CATALOGUE, self.MK) == [(self.MK, "Magic Kingdom Park")] + + def test_an_ambiguous_name_lists_ids_sorted_by_park_name(self) -> None: + with pytest.raises(SystemExit) as caught: + backfill._resolve(self.CATALOGUE, "Hurricane Harbor") + message = str(caught.value) + hits = [c for c in self.CATALOGUE if "hurricaneharbor" in backfill._normalize(c[1])] + assert len(hits) > 5, "the capture no longer exercises a long ambiguity list" + assert f"matches {len(hits)}" in message + listed = [ln.strip() for ln in message.splitlines() if ln.startswith(" ")] + assert len(listed) == len(hits) + for pid, _pname, _did, dname in hits: + assert any(ln.startswith(pid) for ln in listed), pid + assert dname in message + # Sorted by park name, id first so the id is the copy-pasteable part. + assert listed == sorted(listed, key=lambda ln: (ln.split(" ", 1)[1], ln)) + + def test_an_exact_name_two_parks_share_lists_only_those_two(self) -> None: + with pytest.raises(SystemExit) as caught: + backfill._resolve(self.CATALOGUE, "Disneyland Park") + message = str(caught.value) + assert "matches 2" in message + assert "Hong Kong" not in message + + def test_an_unlisted_uuid_is_passed_to_the_api_to_judge(self) -> None: + unknown = "00000000-1111-2222-3333-444444444444" + assert backfill._resolve(self.CATALOGUE, unknown) == [(unknown, unknown)] + + def test_a_non_ascii_query_matches_nothing_rather_than_everything(self) -> None: + # `_normalize` folds CJK to "" and "" is in every string, so a naive + # implementation lists the entire catalogue as candidates. + with pytest.raises(SystemExit) as caught: + backfill._resolve(self.CATALOGUE, "東京") + assert "no park or destination matching" in str(caught.value) + + +class TestListingTheRealCatalogue: + def test_a_filtered_destination_still_reports_its_real_total(self, capsys) -> None: + assert backfill._print_list(_real_catalogue(), "EPCOT") == 0 + out = capsys.readouterr().out + assert "all 6 parks (1 shown)" in out + assert "47f90d2c-e191-4239-a466-5892ef59a88b" in out + + def test_an_unfiltered_listing_covers_every_park(self, capsys) -> None: + assert backfill._print_list(_real_catalogue(), None) == 0 + out = capsys.readouterr().out + assert sum(1 for c in _real_catalogue() if c[0] in out) == 127 + assert "(1 shown)" not in out, "an unfiltered list has nothing to note" + + +class TestTheCsvContractSharedWithTheJavaScriptSdk: + def test_the_columns_match_the_checked_in_contract(self) -> None: + """`themeparks-backfill` is ONE COMMAND WITH TWO IMPLEMENTATIONS. + + This SDK wrote 41 columns while the JavaScript one wrote 32, with + `inParkHoursScheduledMinutes` against `inParkScheduledMinutes`: a customer + using both got two incompatible CSVs of the same park. Both repos hold an + identical copy of this fixture, so a change in one turns red in the other. + """ + contract = json.loads((FIXTURES / "csv_contract.json").read_text(encoding="utf-8")) + assert contract["columns"] == backfill.CSV_COLUMNS + assert backfill._columns_fingerprint("csv") == contract["fingerprint"] + + def test_the_contract_is_not_trivially_satisfiable(self) -> None: + # Guard the oracle: an empty or stub fixture would agree with anything. + contract = json.loads((FIXTURES / "csv_contract.json").read_text(encoding="utf-8")) + assert len(contract["columns"]) == 41 + assert len(contract["fingerprint"]) == 16 + assert contract["columns"][:5] == backfill.IDENTITY_COLUMNS diff --git a/tests/unit/test_history_paging.py b/tests/unit/test_history_paging.py new file mode 100644 index 0000000..a0a764f --- /dev/null +++ b/tests/unit/test_history_paging.py @@ -0,0 +1,158 @@ +"""Paging over daily history, against two real consecutive pages. + +`mk_park_daily_page1.json` and `mk_park_daily_page2.json` are one real request +and the `next` it handed back, captured verbatim. They are the oracle for the +one thing a resumable download depends on: where a page ends and where the +server says to carry on. Page one covers through 2026-08-31, two of its three +entities stop reporting on 2026-08-30, and the server says continue at +2026-09-01 -- so the newest row, the page's last day, and the resume point are +three different values, and a test cannot confuse them by accident. +""" + +from __future__ import annotations + +import json +from pathlib import Path +from typing import Any + +from themeparks._ergonomic.history import HistoryApi, HistoryPage +from themeparks._generated.models import HistoryDailyRow +from themeparks._models_base import ApiModel +from themeparks._raw import _parse_daily_history + +FIXTURES = Path(__file__).resolve().parents[1] / "fixtures" + + +def _page(name: str) -> dict[str, Any]: + return json.loads((FIXTURES / name).read_text(encoding="utf-8")) + + +class _Raw: + """The generated client, answering the two captured pages in order.""" + + def __init__(self, pages: list[dict[str, Any]]) -> None: + self._pages = pages + self.paths: list[str] = [] + + def get_entity_history_daily(self, entity_id: str, start=None, end=None): + self.paths.append(f"first:{start}:{end}") + return _parse_daily_history(self._pages[0]) + + def get_path(self, url: str): + self.paths.append(url) + return self._pages[len(self.paths) - 1] + + +def _api() -> tuple[HistoryApi, _Raw]: + raw = _Raw([_page("mk_park_daily_page1.json"), _page("mk_park_daily_page2.json")]) + return HistoryApi(raw, "75ea578a-adc8-4116-a54d-dccb60765ef9"), raw + + +class TestThePageHook: + def test_reports_the_range_and_next_the_server_gave(self) -> None: + api, _ = _api() + pages: list[HistoryPage] = [] + list(api.days_with_entities("2026-08-01", "2026-09-20", on_page=pages.append)) + assert pages == [ + HistoryPage("2026-08-01", "2026-08-31", _page("mk_park_daily_page1.json")["next"]), + HistoryPage("2026-09-01", "2026-09-20", None), + ] + + def test_fires_only_after_every_row_of_its_page(self) -> None: + # A resumable download checkpoints on this. Firing first would record a + # checkpoint past rows that were never written, and those rows would be + # missing from the file for good -- the one failure worse than + # duplicating them. + api, _ = _api() + order: list[str] = [] + for _ref, row in api.days_with_entities( + "2026-08-01", "2026-09-20", on_page=lambda p: order.append(f"page:{p.end}") + ): + order.append(f"row:{row.date.isoformat()}") + boundary = order.index("page:2026-08-31") + page1 = _page("mk_park_daily_page1.json") + page1_rows = sum(len(e["days"]) for e in page1["entities"]) + # COUNTED, not indexed. `order[boundary - 1]` was the first assertion + # here, and with the hook firing first the boundary lands at index 0 and + # `order[-1]` is the last row of the run, so it passed on the broken + # ordering -- the assertion tested Python's negative indexing, not the + # code. + assert boundary > 0, "the boundary fired before any row" + before = [item for item in order[:boundary] if item.startswith("row:")] + assert len(before) == page1_rows == 64 + page1_dates = {day["date"] for e in page1["entities"] for day in e["days"]} + assert all(item[4:] in page1_dates for item in before) + assert boundary < len(order) - 1, "the second page never arrived" + + def test_the_newest_row_is_not_the_page_boundary(self) -> None: + # The defect this whole mechanism exists for, stated as data: if these + # were equal, resuming from the newest row would be harmless and none of + # this would be needed. + page1 = _page("mk_park_daily_page1.json") + newest = max(day["date"] for e in page1["entities"] for day in e["days"]) + assert page1["range"]["to"] == newest + stops_early = [ + e["name"] for e in page1["entities"] if e["days"][-1]["date"] < page1["range"]["to"] + ] + assert len(stops_early) == 2, stops_early + assert "from=2026-09-01" in page1["next"] + + +class TestRowsKeepEverythingTheApiSent: + def test_the_undocumented_fields_survive_parsing(self) -> None: + # Pydantic drops what the model does not declare, so a stale vendored + # schema silently deletes data on the way in -- not just from the CSV, + # from every Python caller. `unknownMinutes`, `inParkHours` and + # `extremeWaits` were on every row the API returned and in none of the + # models, so they never reached anyone. + api, _ = _api() + rows = [row for _ref, row in api.days_with_entities("2026-08-01", "2026-09-20")] + assert len(rows) == 64 + 44 + assert any(r.unknownMinutes is not None for r in rows) + with_hours = [r for r in rows if r.inParkHours is not None] + assert with_hours, "inParkHours never survived the parse" + assert with_hours[0].inParkHours.scheduledMinutes is not None + + def test_each_row_is_labelled_from_the_response(self) -> None: + api, _ = _api() + page1 = _page("mk_park_daily_page1.json") + names = {e["id"]: (e["name"], e["entityType"]) for e in page1["entities"]} + seen = set() + for ref, _row in api.days_with_entities("2026-08-01", "2026-09-20"): + assert (ref.name, ref.entity_type) == names[ref.id] + seen.add(ref.entity_type) + assert len(seen) == 3, seen + + +class TestAFieldTheSchemaDoesNotKnowSurvives: + """The SDK must not delete data because its vendored spec is a week behind. + + Pydantic's default is to drop undeclared fields. The spec trailed the API by + days, and in that window `unknownMinutes`, `inParkHours` and `extremeWaits` + were deleted at parse time for every caller of `days()` -- not untyped, gone, + with nothing failing and nothing warning. The models inherit + `themeparks._models_base.ApiModel` now, whose only job is `extra="allow"`. + """ + + def test_an_unknown_field_reaches_the_caller(self) -> None: + page = _page("mk_park_daily_page1.json") + page["next"] = None # one page: this is about parsing, not paging + # A field this SDK has never heard of, in the shape a new API field arrives + # in: present on the row, absent from the schema. + for entity in page["entities"]: + for day in entity["days"]: + day["weatherClosureMinutes"] = 41 + + raw = _Raw([page]) + api = HistoryApi(raw, "park") + rows = [row for _ref, row in api.days_with_entities("2026-08-01", "2026-08-31")] + assert rows, "no rows parsed" + # Reachable as an attribute, and in the dump the NDJSON writer uses. + assert rows[0].weatherClosureMinutes == 41 + assert rows[0].model_dump(mode="json")["weatherClosureMinutes"] == 41 + + def test_the_base_class_is_the_reason(self) -> None: + # Pin the mechanism, not just the symptom: a regeneration that loses the + # --base-class flag would put the silent deletion straight back. + assert issubclass(HistoryDailyRow, ApiModel) + assert HistoryDailyRow.model_config.get("extra") == "allow" diff --git a/themeparks/__init__.py b/themeparks/__init__.py index bf65055..cb0bd0f 100644 --- a/themeparks/__init__.py +++ b/themeparks/__init__.py @@ -1,7 +1,12 @@ from themeparks._cache import Cache, CacheConfig, InMemoryLRUCache from themeparks._client import AsyncThemeParks, ThemeParks from themeparks._ergonomic.dates import parse_api_datetime -from themeparks._ergonomic.history import BudgetExhaustedError, HistorySpan +from themeparks._ergonomic.history import ( + BudgetExhaustedError, + EntityRef, + HistoryPage, + HistorySpan, +) from themeparks._ergonomic.live import current_wait_time, iter_queues from themeparks._errors import ( APIError, @@ -17,6 +22,8 @@ "APIError", "AsyncThemeParks", "BudgetExhaustedError", + "EntityRef", + "HistoryPage", "HistorySpan", "RateLimit", "RateLimits", diff --git a/themeparks/_ergonomic/history.py b/themeparks/_ergonomic/history.py index 323510e..8edecb0 100644 --- a/themeparks/_ergonomic/history.py +++ b/themeparks/_ergonomic/history.py @@ -20,7 +20,7 @@ from __future__ import annotations -from collections.abc import AsyncIterator, Iterator +from collections.abc import AsyncIterator, Callable, Iterator from datetime import date as _date from typing import Any, NamedTuple, Union @@ -111,6 +111,28 @@ def _reraise_if_too_long(exc: RateLimitError, max_wait: float) -> None: raise exc +class HistoryPage(NamedTuple): + """One page of daily history, as the server described it. + + `start` and `end` are the park-local days the page ACTUALLY covered, which is + not the range you asked for: a park daily call serves at most 31 days, so a + 50-day request comes back as 31 days plus a `next`. `next_url` is the URL of + the following page, or None on the last one. + + This exists for resumable downloads. A checkpoint taken from the ROWS is + wrong in both directions: the newest row's date can be earlier than the page + covered, because an entity that stopped reporting has no rows for the tail + days, so resuming there re-fetches days already written and duplicates them; + and a half-written page is indistinguishable from a finished one. The page + boundary is the server's own answer to "where do I carry on", so it is the + only safe checkpoint. + """ + + start: str + end: str + next_url: str | None + + class EntityRef(NamedTuple): """Who a history row belongs to, AS THE HISTORY RESPONSE REPORTS IT. @@ -136,6 +158,17 @@ def _ref(entity: Any) -> EntityRef: ) +def _page_of(envelope: DailyEnvelope) -> HistoryPage: + """The page an envelope represents, for :class:`HistoryPage`'s callers.""" + rng = getattr(envelope, "range", None) + nxt = getattr(envelope, "next", None) + return HistoryPage( + getattr(rng, "from_", "") or "", + getattr(rng, "to", "") or "", + nxt or None, + ) + + def _daily_entity_rows(envelope: DailyEnvelope) -> Iterator[tuple[EntityRef, HistoryDailyRow]]: """Yield (entity ref, row), keeping the name the response gave. @@ -197,6 +230,7 @@ def days_with_entities( end: str | _date | None = None, *, max_wait: float = DEFAULT_MAX_WAIT_SECONDS, + on_page: Callable[[HistoryPage], None] | None = None, ) -> Iterator[tuple[EntityRef, HistoryDailyRow]]: """`days()`, but each row arrives with the entity's name and type. @@ -208,10 +242,17 @@ def days_with_entities( It also saves a request: the name is already in the payload, so nothing needs to ask what an id refers to. + + `on_page` is called once every row of a page has been yielded, with a + :class:`HistoryPage`. Checkpoint on that, never on the last row you saw. """ envelope: DailyEnvelope | None = self._first_daily(start, end, max_wait) while envelope is not None: yield from _daily_entity_rows(envelope) + # AFTER the rows, never before: a caller checkpointing on this has to + # be able to trust that everything the page held is already written. + if on_page is not None: + on_page(_page_of(envelope)) envelope = self._next_daily(envelope, max_wait) def days( @@ -220,6 +261,7 @@ def days( end: str | _date | None = None, *, max_wait: float = DEFAULT_MAX_WAIT_SECONDS, + on_page: Callable[[HistoryPage], None] | None = None, ) -> Iterator[tuple[str, HistoryDailyRow]]: """One summary row per park-local day, as (entity id, row). @@ -232,6 +274,8 @@ def days( envelope: DailyEnvelope | None = self._first_daily(start, end, max_wait) while envelope is not None: yield from _daily_rows(envelope) + if on_page is not None: + on_page(_page_of(envelope)) envelope = self._next_daily(envelope, max_wait) def _first_daily( @@ -300,6 +344,7 @@ async def days( end: str | _date | None = None, *, max_wait: float = DEFAULT_MAX_WAIT_SECONDS, + on_page: Callable[[HistoryPage], None] | None = None, ) -> AsyncIterator[tuple[str, HistoryDailyRow]]: try: envelope: DailyEnvelope | None = await self._raw.get_entity_history_daily( @@ -311,6 +356,8 @@ async def days( while envelope is not None: for pair in _daily_rows(envelope): yield pair + if on_page is not None: + on_page(_page_of(envelope)) nxt = getattr(envelope, "next", None) if not nxt: return diff --git a/themeparks/_generated/models.py b/themeparks/_generated/models.py index 868d3b0..f6545eb 100644 --- a/themeparks/_generated/models.py +++ b/themeparks/_generated/models.py @@ -7,7 +7,53 @@ from enum import Enum, IntEnum from typing import Annotated -from pydantic import AwareDatetime, BaseModel, ConfigDict, Field +from pydantic import AwareDatetime, ConfigDict, Field + +from themeparks._models_base import ApiModel + + +class Success(Enum): + """ + Always false. + """ + + boolean_False = False + + +class Type(Enum): + Authentication_failed = "Authentication failed" + + +class Code(IntEnum): + """ + Repeats the HTTP status. + """ + + integer_401 = 401 + + +class Error(ApiModel): + type: Type + message: str + """ + Says whether no credential was sent or the one sent was not accepted. + """ + code: Code | None = None + """ + Repeats the HTTP status. + """ + + +class AuthenticationRequired(ApiModel): + """ + 401: this endpoint needs an API key (X-API-Key header) or a session token, and none was sent or the one sent was not accepted. A revoked, mistyped or truncated key answers this too. + """ + + success: Success + """ + Always false. + """ + error: Error class BoardingGroupState(Enum): @@ -20,7 +66,7 @@ class BoardingGroupState(Enum): CLOSED = "CLOSED" -class DiningAvailability(BaseModel): +class DiningAvailability(ApiModel): partySize: float | None = None """ Available party size @@ -45,19 +91,11 @@ class AttractionType(Enum): OTHER = "OTHER" -class Success(Enum): - """ - Always false. Branch on this rather than on the status alone. - """ - - boolean_False = False - - -class Type(Enum): +class Type1(Enum): Bad_request = "Bad request" -class Code(IntEnum): +class Code1(IntEnum): """ Repeats the HTTP status. """ @@ -65,31 +103,31 @@ class Code(IntEnum): integer_400 = 400 -class Error(BaseModel): - type: Type +class Error1(ApiModel): + type: Type1 message: str """ Says which parameter was rejected and what it must look like, e.g. "Month must be a two-digit number (01-12)". """ - code: Code | None = None + code: Code1 | None = None """ Repeats the HTTP status. """ -class EntityInvalidParameter(BaseModel): +class EntityInvalidParameter(ApiModel): """ - 400: a path parameter is the wrong shape. Checked before the entity is looked up, so a bad `year` or `month` answers 400 whether or not the id exists. + 400: a path parameter is the wrong shape, such as a one-digit `month` or a `year` outside the accepted range. An unknown id returns 404 first. """ success: Success """ - Always false. Branch on this rather than on the status alone. + Always false. """ - error: Error + error: Error1 -class EntityLocation(BaseModel): +class EntityLocation(ApiModel): latitude: float | None = None """ Latitude coordinate of the entity location @@ -100,11 +138,11 @@ class EntityLocation(BaseModel): """ -class Type1(Enum): +class Type2(Enum): Not_found = "Not found" -class Code1(IntEnum): +class Code2(IntEnum): """ Repeats the HTTP status. """ @@ -112,28 +150,28 @@ class Code1(IntEnum): integer_404 = 404 -class Error1(BaseModel): - type: Type1 +class Error2(ApiModel): + type: Type2 message: str """ Names the id that did not resolve, e.g. "Entity 00000000-0000-0000-0000-000000000000 not found". """ - code: Code1 | None = None + code: Code2 | None = None """ Repeats the HTTP status. """ -class EntityNotFound(BaseModel): +class EntityNotFound(ApiModel): """ - 404: nothing resolved from `id`. An id is a UUID, a slug, or — for a destination — its upstream external id; an id malformed enough that the lookup itself rejects it answers 404 as well, rather than 400 or 500. The body is the same whether the entity never existed or has been removed, so it cannot be used to tell those apart. + 404: nothing matches `id` (a UUID, a slug, or a destination's externalId; a malformed id is 404 too). The answer is the same whether it never existed or was removed. """ success: Success """ - Always false. Branch on this rather than on the status alone. + Always false. """ - error: Error1 + error: Error2 class EntityType(Enum): @@ -149,14 +187,14 @@ class EntityType(Enum): SHOW = "SHOW" -class HistoryCoverage(BaseModel): +class HistoryCoverage(ApiModel): firstRecordedAt: date_aliased | None = None """ - First park-local day with recorded history for this entity, or null when nothing has been archived yet. Per-kind detail and gaps: GET /v1/entity/{id}/history/coverage. + First park-local day we hold anything for this entity, or null. Per-field detail: GET /v1/entity/{id}/history/coverage. """ -class HistoryCoverageKindSpan(BaseModel): +class HistoryCoverageKindSpan(ApiModel): """ One live-data field's recorded span for an entity, both ends inclusive. """ @@ -167,22 +205,37 @@ class HistoryCoverageKindSpan(BaseModel): """ last: date_aliased """ - Newest park-local day this field appears in the ARCHIVE. This is not the same as the last day the field was reported, and it does NOT mean the field has stopped: the archive is written two to three days behind live data, so a field being published right now still has a `last` a few days in the past. Every active field looks the same as a withdrawn one here. Use this to know how far back the archive goes and how current it is, not to decide whether a field is still live — GET /v1/entity/{id}/live answers that directly. + The newest day we hold this field for. The coverage call says how far it trails today. """ -class HistoryDailyStats(BaseModel): +class HistoryDailyExtremeWaits(ApiModel): + """ + How many wait readings of 480 minutes or more the day's statistics include. Such readings are usually feed errors (values like 999); they are counted here as a flag and stay in every statistic. Counted per reading while OPERATING. Present only when there were some, and then with both counts. + """ + + standby: int """ - Wait statistics for one park-local day. The percentiles and the mean are weighted by the MINUTES the wait was posted rather than by the number of readings, so a wait that stood for three hours counts three hours and a brief flap does not drag the median; they are sampled at minute resolution and the percentiles are nearest-rank, never interpolated. `min` and `max` are TRUE extremes over every value posted, so a spike too short to be sampled still shows there. Only periods where the entity was OPERATING and published a numeric wait count at all. The block is absent when it never did. + Standby readings of 480 minutes or more included in the day's standby statistics. + """ + singleRider: int + """ + Single-rider readings of 480 minutes or more included in the day's single-rider statistics. + """ + + +class HistoryDailyStats(ApiModel): + """ + Wait statistics for the day. p50, mean and p90 are weighted by how many minutes each wait was showing; min and max are the extremes of every value posted. Only minutes where the entity was OPERATING (not unknown) and the wait was valid count. A wait is valid during the run and on the day it was posted; the wait showing when a ride opens counts from the opening if it last changed within 24 hours; a wait carried over midnight does not count until it changes. Nothing the park reported is excluded: a reading of 480 minutes or more stays in these statistics and is counted in the row's `extremeWaits`. Absent when no minute counted. """ min: int """ - Lowest wait, in minutes, the entity published while OPERATING that day. A true extreme over every value posted, including one that stood for less than a minute — so unlike the percentiles below it is not minute-weighted. + Lowest wait posted during the counted minutes, including one that showed for under a minute. """ p50: int """ - Median wait, nearest-rank over the minute weights (a value actually posted, never interpolated). + Median wait, weighted by minutes; always a value that was actually posted. """ mean: int """ @@ -190,114 +243,114 @@ class HistoryDailyStats(BaseModel): """ p90: int """ - 90th-percentile wait, nearest-rank over the minute weights. + 90th-percentile wait, weighted by minutes. """ max: int """ - Highest wait, in minutes, the entity published while OPERATING that day. A true extreme, as min is: a spike that lasted forty seconds counts here and is invisible to the percentiles, which is the intended difference between the two halves of this block — extremes answer "what did it ever reach", percentiles answer "what was it usually like". + Highest wait posted during the counted minutes, including a brief spike. It can be a feed error: the row's extremeWaits counts readings of 480 minutes or more. """ -class Type2(Enum): +class Type3(Enum): HISTORY_BACKEND_UNAVAILABLE = "HISTORY_BACKEND_UNAVAILABLE" -class Error2(BaseModel): - type: Type2 +class Error3(ApiModel): + type: Type3 message: str -class HistoryErrorBackendUnavailable(BaseModel): +class HistoryErrorBackendUnavailable(ApiModel): """ - 502: the range needs archived history and that backend is temporarily unavailable. The request is retryable. + 502: history is temporarily unavailable. Retry shortly. """ - error: Error2 + error: Error3 -class Type3(Enum): +class Type4(Enum): INVALID_DATE = "INVALID_DATE" -class Error3(BaseModel): - type: Type3 +class Error4(ApiModel): + type: Type4 message: str """ e.g. "from must be a calendar day (YYYY-MM-DD) or an RFC 3339 instant with an offset (e.g. 2026-09-13T14:00:00Z)." """ -class HistoryErrorInvalidDate(BaseModel): +class HistoryErrorInvalidDate(ApiModel): """ 400: a date parameter is not a calendar day or an RFC 3339 instant with an explicit offset, or date was combined with from/to, or from and to mix the two forms, or to was given without from. """ - error: Error3 + error: Error4 -class Type4(Enum): +class Type5(Enum): INVALID_RANGE = "INVALID_RANGE" -class Error4(BaseModel): - type: Type4 +class Error5(ApiModel): + type: Type5 message: str """ e.g. "to must not be before from." """ -class HistoryErrorInvalidRange(BaseModel): +class HistoryErrorInvalidRange(ApiModel): """ 400: both ends parsed, but the range runs backwards (days: to before from; instants: to not after from). """ - error: Error4 + error: Error5 -class Type5(Enum): +class Type6(Enum): NOT_FOUND = "NOT_FOUND" -class Error5(BaseModel): - type: Type5 +class Error6(ApiModel): + type: Type6 message: str -class HistoryErrorNotFound(BaseModel): +class HistoryErrorNotFound(ApiModel): """ 404: no entity with that id. """ - error: Error5 + error: Error6 -class Type6(Enum): +class Type7(Enum): RANGE_TOO_LONG = "RANGE_TOO_LONG" -class Error6(BaseModel): - type: Type6 +class Error7(ApiModel): + type: Type7 message: str """ - Names the cap that was exceeded and the span that was asked for, e.g. "A history call covers at most 31 park-local days (2026-01-01 to 2026-03-01 is 60). Ask for a shorter range." The number is the cap for the path that answered, not a constant: read it from the message rather than hard-coding 31. + Names the limit that applied and your span, e.g. "A history call covers at most 31 park-local days (2026-01-01 to 2026-03-01 is 60)." """ -class HistoryErrorRangeTooLong(BaseModel): +class HistoryErrorRangeTooLong(ApiModel): """ - 400: the range is longer than the path allows. The cap is NOT the same on every path, and this one error type is returned by all of them: GET /v1/entity/{id}/history allows 31 park-local days for a single entity and 1 for a PARK; GET /v1/entity/{id}/history/daily allows 3660 (ten years) for a single entity, and for a PARK serves 31 days a page and gives you `next` for the rest. The message names the cap that applied. Split the ask into consecutive calls. + 400: the range is too long. The limit is not the same on every path: 31 days on /history (1 for a park), and 3660 on /history/daily (a park pages at 31 instead). Split the range into shorter calls. """ - error: Error6 + error: Error7 -class Type7(Enum): +class Type8(Enum): HISTORY_RATE_LIMITED = "HISTORY_RATE_LIMITED" -class Error7(BaseModel): - type: Type7 +class Error8(ApiModel): + type: Type8 message: str """ e.g. "This key can make 600 history requests an hour." @@ -308,23 +361,23 @@ class Error7(BaseModel): """ -class HistoryErrorRateLimited(BaseModel): +class HistoryErrorRateLimited(ApiModel): """ - 429: the caller's hourly history request budget is spent. The budget is separate from the per-minute REST limit and is published per tier in GET /tiers as limits.historyRequestsPerHour (anonymous.historyRequestsPerHour for keyless calls). + 429: your hourly history budget is spent. It is separate from the per-minute limit; the limits per plan are on the pricing page. """ - error: Error7 + error: Error8 -class Type8(Enum): +class Type9(Enum): HISTORY_WINDOW_EXCEEDED = "HISTORY_WINDOW_EXCEEDED" -class Error8(BaseModel): - type: Type8 +class Error9(ApiModel): + type: Type9 message: str """ - A fact and a date, naming no plan and selling nothing. e.g. "This key can see history back to 2026-09-08 (7 days).", or "Requests without an API key can see history back to 2026-09-08 (7 days)." when you sent no key. + e.g. "This key can see history back to 2026-09-08 (7 days)." """ earliestAllowedDate: date_aliased """ @@ -332,13 +385,31 @@ class Error8(BaseModel): """ -class HistoryErrorWindowExceeded(BaseModel): - error: Error8 +class HistoryErrorWindowExceeded(ApiModel): + error: Error9 + + +class Degraded(Enum): + """ + Present, and true, when we could not look far enough back for this response, so the opening may be missing a field we hold. Ask again in a minute for the full opening. + """ + boolean_True = True -class HistoryParkCoverageDepth(BaseModel): + +class DegradedReason(Enum): """ - How far back each entity's record of one field goes, as counts. Calendar years, measured from the day the document was built (`summary.measuredOn`). Counts, not a percentage or an average: a single figure for a park would hide that most of one park's standby entities have two to four years while its oldest reach back to the start of the archive. + Why the lookup was cut short: `timeout`, `error`, or `capacity` when it needed more reading than one request is allowed. + """ + + timeout = "timeout" + error = "error" + capacity = "capacity" + + +class HistoryParkCoverageDepth(ApiModel): + """ + How far back each entity's record of this field goes, as counts per band of calendar years, measured from `summary.measuredOn`. """ fourYearsPlus: int @@ -355,13 +426,13 @@ class HistoryParkCoverageDepth(BaseModel): """ underOneYear: int """ - Entities with under a calendar year, typically something that opened recently rather than a gap. + Entities with under a calendar year, usually something that opened recently. """ -class HistoryParkCoverageEntity(BaseModel): +class HistoryParkCoverageEntity(ApiModel): """ - One entity of the park that history is held for. An entity nothing is held for is ABSENT rather than listed as empty: it is usually a parade, a show or a land, which never had a queue to record, and listing it would read as a gap. + One entity we hold history for. Entities with none (often parades, shows and lands) are absent. """ id: str @@ -377,17 +448,17 @@ class HistoryParkCoverageEntity(BaseModel): """ fields: list[str] """ - The live-data field paths held for this entity, named exactly as the single-entity coverage document names them, so one client type reads both. + Field paths held, named as in the single-entity coverage document. """ stillListed: bool """ - False when the park no longer lists this entity. Its history is still held and still retrievable, and every count in this document includes it; this says where the entity is now, not what the archive has. + false when the park no longer lists this entity. Its history is still available. """ -class HistoryParkCoverageField(BaseModel): +class HistoryParkCoverageField(ApiModel): """ - What one live-data field looks like across the whole park. `entities` counts what is HELD and is never a fraction: entities that have never reported this field are simply absent from the count, because a denominator drawn from entityType would publish parades, shows and lands as missing wait times. + One field across the park. `entities` counts the entities we hold it for; entities that never reported it are not counted. """ entities: int @@ -400,23 +471,23 @@ class HistoryParkCoverageField(BaseModel): """ newest: date_aliased """ - The newest park-local day any entity in the park reported it. A day in the past is not staleness: when a park stops publishing a field the ending is recorded, so the span genuinely stops there. + Newest day any entity reported it. A past day can mean the park stopped publishing the field. """ depth: HistoryParkCoverageDepth -class HistoryParkCoverageSummary(BaseModel): +class HistoryParkCoverageSummary(ApiModel): """ - The park in four numbers. Every one describes what is held; none is a fraction of a total, and nothing here asserts that anything is missing. + Summary figures for the park, all describing what we hold. """ entitiesWithData: int """ - Entities of this park any live-data history is held for - attractions, restaurants, shows and anything else that has ever reported. Larger than the number with wait times: `fields` breaks it down. Counts entities the park no longer lists as well, since their history is still held; those carry `stillListed: false` in `entities`. + Entities we hold any history for, including ones the park no longer lists (`stillListed: false`). """ archiveFrom: date_aliased """ - The earliest park-local day anything in this park was recorded, or null when nothing has been. + Earliest day anything in this park was recorded, or null. """ recordedTo: date_aliased """ @@ -424,40 +495,40 @@ class HistoryParkCoverageSummary(BaseModel): """ retrievableThrough: date_aliased """ - The newest park-local day a caller can actually retrieve. Runs ahead of `recordedTo` by a day or two: the most recent days are served from live data before they are sealed into the archive. Null when nothing is recorded. + The newest day you can ask for, usually today. Null when nothing is recorded. """ measuredOn: date_aliased """ - The park-local day these figures were computed. They move as the archive grows, so a reader comparing two copies of this document needs to know which day each was built. + The day these figures were computed; they grow over time. """ -class HistoryRange(BaseModel): +class HistoryRange(ApiModel): from_: Annotated[str, Field(alias="from")] """ - The requested start. A park-local day comes back verbatim (YYYY-MM-DD); an instant comes back NORMALISED to UTC whole seconds (2026-09-13T14:00:00Z), so an offset or sub-second precision you sent is not echoed back. On a day-granular endpoint such as /history/daily this is ALWAYS a park-local day, even when you asked with an instant: that endpoint's rows are whole days and cannot be sliced finer, so echoing your instant back would claim a precision the data does not have. + The start you asked for. A day comes back as you sent it; an instant comes back normalised to UTC whole seconds. On /history/daily it is always a day. """ to: str """ - The requested end, in the same form as from, and normalised the same way. Omitted instants default to now; omitted days default to today, park-local. The same day-granular rule as from applies on /history/daily. + The end, in the same form as from. Defaults to today (days) or now (instants). """ -class StandbyQueue(BaseModel): +class StandbyQueue(ApiModel): waitTime: float | None = None """ Current standby wait time in minutes """ -class SingleRiderQueue(BaseModel): +class SingleRiderQueue(ApiModel): waitTime: float | None = None """ Current single rider wait time in minutes """ -class BoardingGroupQueue(BaseModel): +class BoardingGroupQueue(ApiModel): allocationStatus: BoardingGroupState | None = None currentGroupStart: float | None = None """ @@ -477,14 +548,14 @@ class BoardingGroupQueue(BaseModel): """ -class PaidStandbyQueue(BaseModel): +class PaidStandbyQueue(ApiModel): waitTime: float | None = None """ Current paid standby wait time in minutes """ -class LiveShowTime(BaseModel): +class LiveShowTime(ApiModel): type: str """ Type of show time entry @@ -501,7 +572,7 @@ class LiveShowTime(BaseModel): class LiveStatusType(Enum): """ - Current operating status of an entity + An entity's status. OPERATING: open and running. DOWN: stopped for now, for example by a breakdown, when it would otherwise be running. CLOSED: not open. REFURBISHMENT: closed for a longer period of maintenance or rebuilding. """ OPERATING = "OPERATING" @@ -510,7 +581,7 @@ class LiveStatusType(Enum): REFURBISHMENT = "REFURBISHMENT" -class Park(BaseModel): +class Park(ApiModel): id: str """ Unique identifier of the park @@ -521,7 +592,34 @@ class Park(BaseModel): """ -class PriceData(BaseModel): +class Success3(Enum): + boolean_False = False + + +class Type10(Enum): + Server_error = "Server error" + + +class Code3(IntEnum): + integer_503 = 503 + + +class Error10(ApiModel): + type: Type10 + message: str + code: Code3 | None = None + + +class PlanUnavailable(ApiModel): + """ + 503: your plan could not be read just now. Try again after the Retry-After header's number of seconds. + """ + + success: Success3 + error: Error10 + + +class PriceData(ApiModel): amount: float | None = None """ Numerical price amount, in the currency's lowest denomination (e.g. cents). null when the item costs money but the provider does not publish an amount; 0 means genuinely free @@ -536,6 +634,30 @@ class PriceData(BaseModel): """ +class Error11(Enum): + Too_Many_Requests = "Too Many Requests" + + +class RateLimited(ApiModel): + """ + 429 from the request limit every call counts against. See "Limits" at the top of this document. + """ + + error: Error11 + message: str + """ + What happened, e.g. "Too Many Requests". + """ + retryAfter: int + """ + Seconds to wait before retrying. Also sent as the Retry-After header. + """ + useInstead: str | None = None + """ + Present when you were looking entities up one guessed slug at a time: the call to make instead, e.g. "GET /v1/destinations". + """ + + class ReturnTimeState(Enum): """ State of return time availability @@ -546,7 +668,7 @@ class ReturnTimeState(Enum): FINISHED = "FINISHED" -class Type9(Enum): +class Type11(Enum): """ Type of schedule entry """ @@ -568,7 +690,34 @@ class SchedulePriceType(Enum): ATTRACTION = "ATTRACTION" -class DestinationEntry(BaseModel): +class V1MeBudget(ApiModel): + """ + One budget: the same figures the RateLimit-* and RateLimit-History-* headers carry. + """ + + unmetered: bool + """ + true when this budget is not metered on your plan. The other fields are then absent, and no RateLimit headers are sent for it. + """ + limit: int | None = None + """ + Requests allowed per window. + """ + windowSeconds: int | None = None + """ + Window length in seconds. + """ + remaining: int | None = None + """ + Requests left in the current window; null if it could not be read just now. + """ + reset: int | None = None + """ + Seconds until the window resets; null if nothing has been counted in this window yet, or if it could not be read. + """ + + +class DestinationEntry(ApiModel): id: str """ Unique identifier of the destination @@ -583,7 +732,7 @@ class DestinationEntry(BaseModel): """ externalId: str | None = None """ - External entity ID from the source data provider + The park operator's own id for this destination """ parks: list[Park] """ @@ -591,14 +740,14 @@ class DestinationEntry(BaseModel): """ -class DestinationsResponse(BaseModel): +class DestinationsResponse(ApiModel): destinations: list[DestinationEntry] """ Array of all destinations """ -class EntityChild(BaseModel): +class EntityChild(ApiModel): id: str """ Unique entity identifier @@ -623,7 +772,7 @@ class EntityChild(BaseModel): """ -class EntityChildrenResponse(BaseModel): +class EntityChildrenResponse(ApiModel): id: str | None = None """ Parent entity identifier @@ -640,9 +789,9 @@ class EntityChildrenResponse(BaseModel): children: list[EntityChild] | None = None -class EntityData(BaseModel): +class EntityData(ApiModel): """ - A single entity. Beyond the properties listed here, an entity may carry additional tag-derived properties named after the tag's slug, for example `minimumHeight` (integer, centimetres) or `mayGetWet` (boolean). The set is open-ended and driven by data rather than fixed by this contract, so clients should read them defensively rather than assume any particular tag is present. + A single entity. It may also carry attribute keys not listed here, named after the attribute, e.g. `minimumHeight` (integer, centimetres) or `mayGetWet` (boolean). Which ones appear varies by entity, so read the ones you need and do not assume any is present. """ model_config = ConfigDict( @@ -680,7 +829,7 @@ class EntityData(BaseModel): location: EntityLocation | None = None externalId: str | None = None """ - Identifier used by the source data provider. + The park operator's own id for this entity. """ slug: str | None = None """ @@ -688,9 +837,9 @@ class EntityData(BaseModel): """ -class HistoryCoverageDocument(BaseModel): +class HistoryCoverageDocument(ApiModel): """ - What history is actually held for one entity, broken down per live-data field. Two different questions, answered separately: `firstRecordedAt`/`lastRecordedAt` and the per-field spans describe the ARCHIVE, while `retrievableThrough` is the newest day a history call could return for this entity — which is normally today, and normally two to three days AHEAD of `lastRecordedAt`. An entity with nothing recorded is a 200 with kinds: {} and every day null, never a 404: "we hold nothing for this entity" is a real, actionable answer and a different claim from "this entity does not exist". + What history we hold for one entity, per field. An entity with nothing recorded is a 200 with `kinds: {}` and null days, not 404. For a PARK the same call returns HistoryParkCoverageDocument. """ id: str @@ -704,25 +853,51 @@ class HistoryCoverageDocument(BaseModel): """ firstRecordedAt: date_aliased | None = None """ - First park-local day with any recorded history for this entity, or null when nothing has been archived yet. May legitimately be earlier than any individual kind's first: the two are recorded independently and answer slightly different questions. + First park-local day we hold anything for this entity, or null. """ lastRecordedAt: date_aliased | None = None """ - The newest `last` across `kinds` — the most recent park-local day we hold anything at all for this entity — or null when nothing is recorded. Like the per-field `last`, this tracks the ARCHIVE and lags live data by two to three days, so it sits in the past for an entity reporting normally. + The newest `last` across `kinds`, or null. """ retrievableThrough: date_aliased | None = None """ - The newest park-local day GET /v1/entity/{id}/history and .../history/daily could return data for this entity — what you can ASK FOR, as opposed to what has been filed. Normally TODAY for an entity still reporting, because those endpoints answer from recent readings as well as the archive, and it therefore sits AHEAD of `lastRecordedAt` by two to three days for a healthy entity. That gap is the whole point of this field: `lastRecordedAt` and every per-field `last` describe the ARCHIVE only, and reading them as capability is what makes coverage look as though it has stopped a couple of days short. It is the LATER of `lastRecordedAt` and the newest day recent readings can still answer for — so an entity that stopped reporting long ago reports its archive day here, NOT today, and the field never promises data that is not there. Recent readings do not go back indefinitely, so a day is reported here only if one of those endpoints can actually return it: an entity whose last reading is old enough falls back to its archive day rather than naming the day that reading was taken. Null only when we hold nothing for this entity in either place. A request for a range up to this day can still be narrowed by your tier's history window, which bounds how far BACK you may ask, never how recent. + The newest day you can ask /history and /history/daily for: usually today for an entity still reporting, or its last recorded day for one that stopped long ago. Null when we hold nothing. Your history window limits how far back you can ask. """ kinds: dict[str, HistoryCoverageKindSpan] """ - Keyed by LIVE-DATA PATH (status, queue.StandbyQueue, showtimes, ...), not by the internal kind name, so a key matches straight against what GET /v1/entity/{id}/live and GET /v1/entity/{id}/history return. A kind the entity never reported is ABSENT here and absent from every /history row — there is no zero-span entry for it. See `last` for why a span ending in the past does NOT mean the field stopped being reported: the archive lags live data by two to three days, so an actively published field ends in the past too. Every span here is archive-only; `retrievableThrough` is the entity-wide answer to how recent a day you can actually ask for. + Keyed by live-data path (status, queue.StandbyQueue, showtimes and so on), as /live and /history name them. Fields the entity never reported are left out. """ -class HistoryDailyRow(BaseModel): +class HistoryDailyInParkHours(ApiModel): """ - One park-local day of an entity's history reduced to the numbers a crowd calendar needs. standby, singleRider and showCount are ABSENT rather than null when there is nothing to report: an absent standby means no numeric standby wait was in force while OPERATING for a whole sampled minute that day — the statistics are sampled at minute resolution, so a wait published only inside a sub-minute window produces no block at all, even though /history records it and the day's operatingMinutes count it. The same applies to singleRider. An absent showCount means the entity published no showtimes. showCount counts distinct performance start times in the local day. + The day's numbers limited to the park's published hours. Present only when the park published hours that day. + """ + + scheduledMinutes: int + """ + Minutes of the day inside the park's published hours: its schedule entries of type OPERATING, EXTRA_HOURS and TICKETED_EVENT. Hours from the previous day's schedule that run past midnight are not included. Today counts elapsed minutes only. + """ + operatingMinutes: int + """ + Of scheduledMinutes, the minutes the entity was OPERATING. + """ + downMinutes: int + """ + Of scheduledMinutes, the minutes the entity was DOWN. + """ + unknownMinutes: int + """ + Of scheduledMinutes, the minutes with no known status. + """ + standby: HistoryDailyStats | None = None + singleRider: HistoryDailyStats | None = None + extremeWaits: HistoryDailyExtremeWaits | None = None + + +class HistoryDailyRow(ApiModel): + """ + One day of an entity's history. standby, singleRider, extremeWaits, showCount and inParkHours are absent when there is nothing to report; a standby block needs a valid wait showing for at least one whole minute. A row without unknownMinutes was counted under earlier rules, and also lacks extremeWaits and inParkHours. """ date: date_aliased @@ -731,35 +906,41 @@ class HistoryDailyRow(BaseModel): """ firstOperatingAt: AwareDatetime | None = None """ - UTC instant (whole seconds) the entity first became OPERATING on this park-local day, or null if it never did. A day that opened already OPERATING reports the start of the local day. + When the entity changed to OPERATING that day (UTC, whole seconds). null when it did not change to OPERATING that day; it may still have operated, carried over from the day before. """ lastClosedAt: AwareDatetime | None = None """ - UTC instant (whole seconds) of the last transition out of OPERATING or DOWN into CLOSED or REFURBISHMENT after firstOperatingAt, or null if there was none. For a park closing after local midnight this instant falls on the following UTC day. + When the last run that started this day closed (UTC, whole seconds). A run lasts from a change to OPERATING until the next change to CLOSED or REFURBISHMENT, DOWN periods included. A run still going at midnight closes the next day and is reported on this row. null if it did not close by the end of the next day, or if the close came after 4 hours or more outside published hours with nothing to confirm the run (see unknownMinutes: a change of status on a day with published hours, a change to any field other than showtimes on a day without). """ operatingMinutes: int """ - Minutes of this park-local day the entity was OPERATING. Minutes with no observed state count as neither operating nor down, so the two counters need not add up to the length of the day. For TODAY the count covers only the minutes that have already elapsed, so it grows through the day and is final once the day ends: a caller polling today's row sees it rise, which is the day filling in rather than the answer changing. + Minutes OPERATING with a known status. Today counts elapsed minutes only. CLOSED and REFURBISHMENT minutes are not counted, so operatingMinutes, downMinutes and unknownMinutes need not add up to the day. """ downMinutes: int """ - Minutes of this park-local day the entity was DOWN. As with operatingMinutes, today's count covers only the elapsed part of the day. + Minutes DOWN. Today counts elapsed minutes only. + """ + unknownMinutes: int | None = None + """ + Minutes with no known status: none was reported, or the entity was OPERATING outside the park's published hours with nothing to confirm it for 4 hours or more. On a day the park published hours, only a change of status confirms it, and a status carried over midnight is unknown outside those hours until it changes. On a day with no published hours, or for an entity outside any park, a change to any field other than showtimes (status, a wait or anything else) confirms it, and for a status carried over midnight the hours count from the last change before midnight. Absent on rows counted under earlier rules. """ standby: HistoryDailyStats | None = None singleRider: HistoryDailyStats | None = None + extremeWaits: HistoryDailyExtremeWaits | None = None showCount: int | None = None """ Distinct performance start times whose park-local day is this day. Present only for entities that published showtimes on the day. """ + inParkHours: HistoryDailyInParkHours | None = None changes: int """ - Number of history rows recorded on this day, i.e. instants at which any kind changed. Short flaps are counted as they happened; this is not a cleaned figure. + How many times any field changed this day. """ -class HistoryParkCoverageDocument(BaseModel): +class HistoryParkCoverageDocument(ApiModel): """ - What history is held across a whole PARK. GET /v1/entity/{id}/history/coverage returns this shape when the entity is a PARK; every other entityType, a DESTINATION included, returns HistoryCoverageDocument for that entity alone. A park records nothing itself, so without this rollup the honest-looking answer for a park would be "nothing". It reports depth and breadth only: it does not detect missing days and does not judge whether a recorded value was correct. + What history we hold across a whole PARK, returned by /history/coverage when the entity is a PARK. Its names map to the single-entity document: `fields` is `kinds`, and each `from` and `newest` is a `first` and `last`. It reports depth and breadth; it does not find missing days. """ id: str @@ -769,7 +950,7 @@ class HistoryParkCoverageDocument(BaseModel): destinationId: str | None = None timezone: str """ - IANA timezone the park-local days are resolved in. The park's children inherit it. + IANA timezone of the park. """ summary: HistoryParkCoverageSummary fields: dict[str, HistoryParkCoverageField] @@ -782,9 +963,9 @@ class HistoryParkCoverageDocument(BaseModel): """ -class HistoryParkEntityDaily(BaseModel): +class HistoryParkEntityDaily(ApiModel): """ - One entity of a park in a park DAILY response: its identity, the first day it has any history, and its day rows. The park itself appears as an entry too when it has history of its own (match it by id against the envelope's id). An entity with no history at all is ABSENT from entities[]. + One entity of the park, with its first recorded day and its rows. The park appears too if it has history of its own. """ id: str @@ -793,11 +974,11 @@ class HistoryParkEntityDaily(BaseModel): coverage: HistoryCoverage days: list[HistoryDailyRow] """ - One row per park-local day, ascending by date, exactly as GET /v1/entity/{id}/history/daily returns for this entity on its own. A day with no data is ABSENT, and an entity with history but nothing in the requested days has an empty array rather than vanishing from entities[] — so the entity list keeps its shape from one page to the next. + The entity's rows for this page, as /history/daily returns them for the entity alone. Empty when it has no data in these days. """ -class ReturnTimeQueue(BaseModel): +class ReturnTimeQueue(ApiModel): state: ReturnTimeState | None = None returnStart: AwareDatetime | None = None """ @@ -809,14 +990,18 @@ class ReturnTimeQueue(BaseModel): """ -class PaidReturnTimeQueue(BaseModel): +class PaidReturnTimeQueue(ApiModel): state: ReturnTimeState | None = None returnStart: AwareDatetime | None = None returnEnd: AwareDatetime | None = None price: PriceData | None = None -class LiveQueue(BaseModel): +class LiveQueue(ApiModel): + """ + The queues an entity has, each present only when the entity publishes it. StandbyQueue: the ordinary line. SINGLE_RIDER: a separate line for guests riding alone. RETURN_TIME: a free reservation for a later time window. PAID_RETURN_TIME: the same, paid for. BOARDING_GROUP: a virtual queue that calls groups by number. PAID_STANDBY: a paid line with its own wait. + """ + STANDBY: StandbyQueue | None = None SINGLE_RIDER: SingleRiderQueue | None = None RETURN_TIME: ReturnTimeQueue | None = None @@ -825,7 +1010,7 @@ class LiveQueue(BaseModel): PAID_STANDBY: PaidStandbyQueue | None = None -class SchedulePriceObject(BaseModel): +class SchedulePriceObject(ApiModel): type: SchedulePriceType | None = None """ Type of price object @@ -848,7 +1033,32 @@ class SchedulePriceObject(BaseModel): """ -class EntityLiveData(BaseModel): +class V1Me(ApiModel): + """ + Your plan and what is left of it. `rateLimit` is the request limit every call counts against, this one included. `historyRateLimit` is the hourly budget the history calls count against; reading it here does not spend it. Figures are per account, so every key on one account reports the same. + """ + + tier: str + """ + Your plan, e.g. "free", "pro", "business", "enterprise". + """ + historyDays: int + """ + How many days back your history calls may reach, today included. null when `historyAllArchive` is true. + """ + historyAllArchive: bool + """ + true when this key can read everything we hold for an entity, back to the first day we recorded it. `historyDays` and `historyEarliestDate` are then null. + """ + historyEarliestDate: date_aliased + """ + The earliest day (YYYY-MM-DD) a history call may ask for, computed in UTC. Each park counts days in its own time zone, so near midnight this can be off by one. Earlier days answer 403 HISTORY_WINDOW_EXCEEDED. null when `historyAllArchive` is true. + """ + rateLimit: V1MeBudget + historyRateLimit: V1MeBudget + + +class EntityLiveData(ApiModel): id: str """ Entity identifier @@ -869,7 +1079,7 @@ class EntityLiveData(BaseModel): diningAvailability: list[DiningAvailability] | None = None -class EntityLiveDataResponse(BaseModel): +class EntityLiveDataResponse(ApiModel): id: str | None = None """ Entity identifier @@ -886,9 +1096,9 @@ class EntityLiveDataResponse(BaseModel): liveData: list[EntityLiveData] | None = None -class HistoryDailyEnvelope(BaseModel): +class HistoryDailyEnvelope(ApiModel): """ - A day-by-day summary of one entity's history. GET /v1/entity/{id}/history/daily returns this shape for every entityType EXCEPT PARK; for a PARK the same path returns HistoryParkDailyEnvelope, which carries an entities[] array instead of this envelope's coverage/days block. range.from and range.to ALWAYS echo park-local calendar days (YYYY-MM-DD), even when the call supplied RFC 3339 instants, because this endpoint summarises whole park-local days and an instant echo would claim a precision the rows do not have. TODAY's row is the day so far — its counters cover only elapsed minutes and grow as the day does — and a response to a request without an API key may be up to an hour old, so a poller can see a value that far behind the live feed. If you need the current state of an entity rather than its day so far, GET /v1/entity/{id}/live is not cached that way. + One entity's daily summary. For a PARK the same call returns HistoryParkDailyEnvelope. range is always whole park-local days (YYYY-MM-DD), and today's row is the day so far. """ id: str @@ -904,17 +1114,17 @@ class HistoryDailyEnvelope(BaseModel): coverage: HistoryCoverage days: list[HistoryDailyRow] """ - One row per park-local day, ascending by date. A day with no data is ABSENT: there is no row of zeroes, because "we have nothing for this day" is a different claim from "observed, closed all day". + One row per day, ascending. Days with no data are left out. """ next: str | None = None """ - URL of the next page, or null. Always null for a single entity: a daily call is unpaged. Park calls page; see HistoryParkDailyEnvelope. + Always null for a single entity. """ -class HistoryOpening(BaseModel): +class HistoryOpening(ApiModel): """ - The full live-data state effective at the start of the range, in the same shape as a row. A key is present only when the entity has that kind. When the value is unknown at that instant (typically an older range, answered from the archive rather than from recent readings) the kind carries its EMPTY live value rather than a null container — an unknown standby is {"waitTime": null}, an unknown showtimes list is [] — and status, which has no empty value, is null. + The `opening` object: the complete state at the start of the range, in the same shape as a row. A field is present only if the entity has it. An unknown value is its empty form (standby `{"waitTime": null}`, showtimes `[]`); an unknown status is null. """ time: AwareDatetime @@ -923,13 +1133,19 @@ class HistoryOpening(BaseModel): """ observedAt: AwareDatetime | None = None """ - The UTC instant (whole seconds) at which this state was actually OBSERVED, as opposed to `time`, which is the start of the range you asked for. The two are different questions and only this one tells you whether to trust the state. - - A state carried forward from before the range has an `observedAt` BEFORE `time` — sometimes long before, because the lookup is deliberately unbounded and returns the last reading of each kind at any age. So a day we hold nothing for still reports the newest state we ever saw, which is usually right and occasionally very wrong: a ride whose feed simply stopped mid-operation carries its last wait time forward indefinitely. - - Compare it against `range.from` to decide. Equal to or after the range start means the state was seen inside the range. Before it means carried forward, and how far back you tolerate is yours to choose — a reading an hour before the day began is ordinary, one from three weeks earlier is not evidence about this day. Absent means nothing survives to carry: we can say nothing at all about the state at the start of this range. - - It is the NEWEST instant any kind in this opening was seen. Kinds can be observed at different moments, so an older kind may be staler than this field suggests; treat it as the most generous reading of the opening's age, not a guarantee about every key. Days that genuinely have data are better read from the rows, which carry their own `time`. + When this state was last seen (UTC, whole seconds): the newest entry in `observedAtByKind`. Earlier than `time` means it was carried in from before the range, possibly from long ago; a ride whose feed stopped keeps its last values. Absent means we hold nothing for this entity from before the range, unless `degraded` is set. + """ + observedAtByKind: dict[str, AwareDatetime] | None = None + """ + When each field of the opening was last seen (UTC, whole seconds), keyed by its path: `status`, `queue.StandbyQueue`, `showtimes` and so on. A wait that did not change for weeks is dated weeks back, so check a queue's own entry before treating it as current. Usually, but not always, the moment the value last changed. + """ + degraded: Degraded | None = None + """ + Present, and true, when we could not look far enough back for this response, so the opening may be missing a field we hold. Ask again in a minute for the full opening. + """ + degradedReason: DegradedReason | None = None + """ + Why the lookup was cut short: `timeout`, `error`, or `capacity` when it needed more reading than one request is allowed. """ status: str | None = None """ @@ -939,9 +1155,9 @@ class HistoryOpening(BaseModel): showtimes: list[LiveShowTime] | None = None -class HistoryParkDailyEnvelope(BaseModel): +class HistoryParkDailyEnvelope(ApiModel): """ - A day-by-day summary of a whole PARK: every entity of the park that has history, in one call. GET /v1/entity/{id}/history/daily returns THIS shape when the entity is a PARK (entityType: "PARK") and HistoryDailyEnvelope for every other entityType, so a client should branch on the presence of entities[] or on the entity's type. range.from and range.to are always park-local calendar days, and range.to is THIS PAGE's last day rather than the whole range you asked for: a call serves at most 31 park-local days and next carries the rest. + The daily summary of a whole PARK, returned by /history/daily when the entity's entityType is PARK. A call serves up to 31 park-local days; `range.to` is this page's last day. """ id: str @@ -951,22 +1167,22 @@ class HistoryParkDailyEnvelope(BaseModel): destinationId: str | None = None timezone: str """ - IANA timezone the park-local days are resolved in. The park is the authority on where its day boundaries fall, so every entity below is summarised in THIS zone. + IANA timezone of the park; every entity is summarised in it. """ range: HistoryRange entities: list[HistoryParkEntityDaily] """ - One entry per entity of the park that has history, ascending by name. An entity with no history is ABSENT — never an entry of nulls or an empty days[] — because "we hold nothing for this entity" is a different claim from "we hold nothing for these days". The park itself is included when it has history of its own. + One entry per entity with history, by name. Entities with no history are absent. """ next: str | None = None """ - Absolute URL of the next page, or null on the last one. A park daily call serves at most 31 park-local days and pages by DAY: the next URL repeats your other parameters with from advanced past this page's last day. Entity order never affects paging. + Absolute URL of the next page (your parameters kept, `from` moved on), or null on the last page. Up to 31 park-local days a page. """ -class HistoryRow(BaseModel): +class HistoryRow(ApiModel): """ - One row per instant at which any kind changed. Every present kind is carried forward, so a row is the complete live-data object at that instant (same keys, nesting and enum values as GET /v1/entity/{id}/live). + One row per moment any field changed. Every field is carried forward, so each row is the complete live data at that moment (same keys as GET /v1/entity/{id}/live). """ time: AwareDatetime @@ -982,7 +1198,7 @@ class HistoryRow(BaseModel): showtimes: list[LiveShowTime] | None = None -class ScheduleEntry(BaseModel): +class ScheduleEntry(ApiModel): date: str """ Schedule date @@ -995,7 +1211,7 @@ class ScheduleEntry(BaseModel): """ Closing time """ - type: Type9 + type: Type11 """ Type of schedule entry """ @@ -1009,9 +1225,9 @@ class ScheduleEntry(BaseModel): """ -class HistoryEnvelope(BaseModel): +class HistoryEnvelope(ApiModel): """ - One entity's history as full-state change rows. GET /v1/entity/{id}/history returns this shape for every entityType EXCEPT PARK; for a PARK the same path returns HistoryParkRawEnvelope, which carries an entities[] array instead of this envelope's coverage/opening/history block. + One entity's history. For a PARK the same call returns HistoryParkRawEnvelope instead. """ id: str @@ -1032,13 +1248,13 @@ class HistoryEnvelope(BaseModel): """ next: str | None = None """ - URL of the next page, or null. Always null for a single entity (a call covers up to 31 days). Park calls page; see HistoryParkRawEnvelope. + Always null for a single entity. """ -class HistoryParkEntityRaw(BaseModel): +class HistoryParkEntityRaw(ApiModel): """ - One entity of a park in a park RAW response: the same coverage, opening and history block GET /v1/entity/{id}/history returns for that entity on its own, so one client type reads both. The park itself appears as an entry too when it has history of its own (match it by id against the envelope's id). + One entity of the park, with the same coverage, opening and history as a single-entity call. The park appears too if it has history of its own. """ id: str @@ -1048,13 +1264,13 @@ class HistoryParkEntityRaw(BaseModel): opening: HistoryOpening history: list[HistoryRow] """ - Ascending by time. Empty when this entity recorded no change in the requested range, which is a different claim from having no history at all — an entity with no history is absent from entities[]. + Ascending by time. Empty when nothing changed in the range. """ -class HistoryParkRawEnvelope(BaseModel): +class HistoryParkRawEnvelope(ApiModel): """ - Full-state change rows for a whole PARK: every entity of the park that has history, in one call, for exactly 1 park-local day. GET /v1/entity/{id}/history returns THIS shape when the entity is a PARK (entityType: "PARK") and HistoryEnvelope for every other entityType, so a client should branch on the presence of entities[] or on the entity's type. A range spanning more than one day is 400 RANGE_TOO_LONG: a day of a park is every recorded change for every entity in it, and the one-day limit is what bounds that. Ask day by day. + History of a whole PARK for 1 park-local day, returned by /history when the entity's entityType is PARK. A longer range is 400 RANGE_TOO_LONG. """ id: str @@ -1064,20 +1280,20 @@ class HistoryParkRawEnvelope(BaseModel): destinationId: str | None = None timezone: str """ - IANA timezone the park-local day is resolved in. The park is the authority on where its day boundaries fall, so every entity below is resolved in THIS zone. + IANA timezone of the park; every entity is resolved in it. """ range: HistoryRange entities: list[HistoryParkEntityRaw] """ - One entry per entity of the park that has history, ascending by name. An entity with no history is ABSENT — never an entry of nulls — because "we hold nothing for this entity" is a different claim from "this entity recorded no change today". The park itself is included when it has history of its own. + One entry per entity with history, by name. Entities with no history are absent. """ next: str | None = None """ - Always null: a park history call covers one park-local day, so there is never a next page. + Always null: one day per call. """ -class ParkSchedule(BaseModel): +class ParkSchedule(ApiModel): id: str | None = None """ Entity identifier @@ -1097,7 +1313,7 @@ class ParkSchedule(BaseModel): schedule: list[ScheduleEntry] | None = None -class EntityScheduleResponse(BaseModel): +class EntityScheduleResponse(ApiModel): id: str | None = None """ Entity identifier diff --git a/themeparks/_models_base.py b/themeparks/_models_base.py new file mode 100644 index 0000000..7cb27e6 --- /dev/null +++ b/themeparks/_models_base.py @@ -0,0 +1,29 @@ +"""The base class every generated model inherits. + +Its whole job is one line of configuration: keep fields the schema does not +declare. + +Pydantic's default is to DROP them. The vendored spec lagged the API by a few +days, and in that window `unknownMinutes`, `inParkHours` and `extremeWaits` were +deleted on the way in -- not merely untyped, gone, for every caller of `days()`. +Nothing failed and nothing warned; the data simply was not there. The TypeScript +SDK never had this problem because its types are erased at runtime, so the same +stale schema cost it nothing but autocompletion. + +With `extra="allow"` a field the API adds tomorrow survives parsing, reaches +`model_dump()`, and reaches anyone writing rows to a file, before this SDK knows +it exists. It stops being a deadline. + +Reading an extra field in typed code still needs a regeneration, which is what the +nightly spec-drift job is for. This is about not losing it in the meantime. +""" + +from __future__ import annotations + +from pydantic import BaseModel, ConfigDict + + +class ApiModel(BaseModel): + """A response object: validated on what we know, lossless on what we do not.""" + + model_config = ConfigDict(extra="allow") diff --git a/themeparks/backfill.py b/themeparks/backfill.py index 682dd95..c22187d 100644 --- a/themeparks/backfill.py +++ b/themeparks/backfill.py @@ -47,23 +47,42 @@ import argparse import contextlib -import csv +import hashlib import json import os +import re import sys import unicodedata -from datetime import date +from datetime import date, datetime +from datetime import date as _date from pathlib import Path -from typing import Any, NamedTuple, TextIO, Union - -from themeparks import APIError, BudgetExhaustedError, RateLimitError, ThemeParks -from themeparks._ergonomic.history import EntityRef - -USER_AGENT = "themeparks-backfill/1" +from typing import Any, NamedTuple, TextIO, Union, get_args, get_origin +from urllib.parse import parse_qs, urlparse + +from pydantic import BaseModel + +from themeparks import ( + APIError, + BudgetExhaustedError, + NetworkError, + RateLimitError, + ThemeParks, + ThemeParksError, +) +from themeparks import TimeoutError as ApiTimeoutError # noqa: A004 - the SDK's, not the builtin's +from themeparks._client import PACKAGE_VERSION, _default_user_agent +from themeparks._ergonomic.history import EntityRef, HistoryPage +from themeparks._generated.models import HistoryDailyRow + +# The command's identity IN FRONT OF the SDK's, not instead of it. It used to be +# the literal "themeparks-backfill/1": a hardcoded 1 that could never match a +# release, and it replaced the SDK's user agent entirely, so a support question +# about a bad download had no version to work from at either end. +USER_AGENT = f"themeparks-backfill/{PACKAGE_VERSION} {_default_user_agent()}" EX_TEMPFAIL = 75 -CSV_COLUMNS = [ +IDENTITY_COLUMNS = [ # Identity first. A reader opening this in a spreadsheet should know what a # row is before they reach the numbers, and a table loaded from several files # needs parkId to tell them apart. @@ -72,47 +91,187 @@ "entityId", "entityName", "entityType", - "date", - "firstOperatingAt", - "lastClosedAt", - "operatingMinutes", - "downMinutes", - "showCount", - "changes", - "standbyMin", - "standbyP50", - "standbyMean", - "standbyP90", - "standbyMax", - "singleRiderP50", - "singleRiderMax", ] +def _flatten_columns(model: type[BaseModel], prefix: str = "") -> list[str]: + """Every scalar in a daily row, as one flat column name each. + + DERIVED FROM THE MODEL, not typed out. The hand-written list had drifted three + ways at once: `unknownMinutes` and the whole `inParkHours` block were on every + row the API returns and in no column, `extremeWaits` likewise, and + `singleRider` carried two of its five percentiles while `standby` carried all + five. Ten of thirty-six fields were missing, silently, from a file people pay + for. Generated from the schema it cannot drift again: regenerate the models + and the columns follow. + """ + columns: list[str] = [] + for name, field in model.model_fields.items(): + inner = _stats_model(field.annotation) + if inner is None: + columns.append( + f"{prefix}{name}" if prefix == "" else f"{prefix}{name[0].upper()}{name[1:]}" + ) + continue + head = name if prefix == "" else f"{prefix}{name[0].upper()}{name[1:]}" + columns.extend(_flatten_columns(inner, head)) + return columns + + +#: Origins whose arguments are ELEMENTS, not nested blocks to flatten. +_COLLECTION_ORIGINS = (list, set, frozenset, tuple, dict) + + +def _stats_model(annotation: Any) -> type[BaseModel] | None: + """The nested model an annotation wraps, or None for a scalar. + + Every nested block on a daily row is optional, so the annotation is a union + with None and the model has to be dug out of it. + """ + # A LIST OR DICT OF MODELS IS NOT A NESTED BLOCK. `get_args(list[Stats])` is + # `(Stats,)`, so the model was found and the annotation treated as one object: + # phantom columns in the header, then `AttributeError: type object 'list' has + # no attribute 'model_fields'` on the first row. Any array-of-objects the API + # adds would have done it. + if get_origin(annotation) in _COLLECTION_ORIGINS: + return None + candidates = [annotation, *get_args(annotation)] + for candidate in candidates: + for unwrapped in (candidate, *get_args(candidate)): + if isinstance(unwrapped, type) and issubclass(unwrapped, BaseModel): + return unwrapped + return None + + +DATA_COLUMNS = _flatten_columns(HistoryDailyRow) +CSV_COLUMNS = [*IDENTITY_COLUMNS, *DATA_COLUMNS] + +# The `` rule is not injective: a top-level `standbyMin` alongside +# `standby.min` would produce one name twice, and `DictWriter` then writes the same +# value into both slots under a right-looking header. No collision exists in +# today's schema; this makes the next one a failure at import rather than a wrong +# number in a customer's file. +if len(set(CSV_COLUMNS)) != len(CSV_COLUMNS): + _seen: set[str] = set() + _dupes = sorted({c for c in CSV_COLUMNS if c in _seen or _seen.add(c)}) # type: ignore[func-returns-value] + raise AssertionError(f"the daily row flattens to duplicate column names: {_dupes}") + + +def _cells(value: Any, prefix: str = "") -> dict[str, Any]: + """One model flattened to `{column: value}`, mirroring `_flatten_columns`.""" + out: dict[str, Any] = {} + for name, field in type(value).model_fields.items(): + column = f"{prefix}{name}" if prefix == "" else f"{prefix}{name[0].upper()}{name[1:]}" + item = getattr(value, name, None) + if _stats_model(field.annotation) is not None: + if item is None: + # An absent block is absent for a reason: no wait was in force, or + # the park published no hours. Empty cells, never zeroes -- a zero + # would read as "measured, and it was nothing". + nested = _require(_stats_model(field.annotation)) + for column_name in _flatten_columns(nested, column): + out[column_name] = "" + else: + out.update(_cells(item, column)) + continue + out[column] = _scalar(item) + return out + + +def _require(model: type[BaseModel] | None) -> type[BaseModel]: + if model is None: # pragma: no cover - _cells only calls this when it is not + raise AssertionError("expected a nested model") + return model + + +#: Characters that make a spreadsheet treat a cell as a formula rather than text. +#: Tab and CR are included because Excel strips them and then reads what follows. +_FORMULA_LEADERS = ("=", "+", "-", "@", "\t", "\r") + +#: A number, strictly: no surrounding whitespace, no sign-only, no `\t5`. The +#: JavaScript SDK uses the same pattern. `float()` would call `"\t5"` numeric and +#: leave a tab-led cell undefended in one SDK and defended in the other, for a +#: file the two are supposed to write identically. +_NUMERIC = re.compile(r"^[+-]?(\d+\.?\d*|\.\d+)([eE][+-]?\d+)?$") + + +def _defuse(text: str) -> str: + """Prefix a cell that a spreadsheet would execute rather than display. + + A ride named `=cmd|...` is a formula to Excel. Every string that reaches a cell + here comes from the API -- park and entity names -- and there is no path from a + command-line argument into one, so this needs an upstream park feed to publish + such a name. Cheap enough to do anyway. + + NUMERIC CELLS ARE LEFT ALONE, which is why this is not a bare startswith on the + tuple: `-5` is a number and must stay one. Prefixing it would turn every + negative value in the file into text and break arithmetic in the tool this + exists to protect. + """ + if not text.startswith(_FORMULA_LEADERS) or _NUMERIC.match(text): + return text + return "'" + text + + +def _scalar(value: Any) -> Any: + """A cell a spreadsheet can read: dates and times as ISO, None as empty. + + UTC is written `Z`, not `+00:00`. Both are valid ISO 8601 and mean the same + instant, but `Z` is what the API sends and what the JavaScript SDK's identical + command writes, and a paid export of the same park should not differ by + language. Python's `isoformat()` is the only reason it did. + """ + if value is None: + return "" + if isinstance(value, datetime): + return value.isoformat().replace("+00:00", "Z") + if isinstance(value, _date): + return value.isoformat() + unwrapped = getattr(value, "value", value) + return _defuse(unwrapped) if isinstance(unwrapped, str) else unwrapped + + +def _csv_cell(value: Any) -> str: + """One cell, quoted exactly as the JavaScript SDK quotes it. + + NOT `csv.DictWriter`, and the reason is a real defect. With + `lineterminator="\n"`, the csv module's QUOTE_MINIMAL does not quote a bare + carriage return on Python 3.9 or 3.10 -- it only quotes characters that appear + in the line terminator -- so an entity name containing one produced a row that + parsed as two, with every later column shifted. 3.11 changed the module to + always quote CR and LF, so the bug was invisible on a modern interpreter and + live on two supported ones. + + Relying on stdlib behaviour that moved between 3.10 and 3.11 cannot give a file + that is byte-identical across Python versions, let alone identical to the + JavaScript SDK's. Ten lines of explicit quoting can. + """ + text = "" if value is None else str(value) + if any(ch in text for ch in ('"', ",", "\r", "\n")): + escaped = text.replace('"', '""') + return f'"{escaped}"' + return text + + +def _csv_line(cells: dict[str, Any]) -> str: + """A row, in column order, LF-terminated. + + LF, not RFC 4180's CRLF: every reader accepts either, and this command also + writes NDJSON with LF and has a JavaScript twin that writes LF, so one park + should not come back as three different byte streams. + """ + return ",".join(_csv_cell(cells.get(column)) for column in CSV_COLUMNS) + "\n" + + def _csv_row(ref: EntityRef, row: Any, ident: _RowIdentity) -> dict[str, Any]: - """Flatten the nested standby/singleRider stats into one wide row.""" - standby = row.standby - single = row.singleRider + """One CSV row: the run's identity, then every field of the day's row.""" return { "parkId": ident.park.id, - "parkName": ident.park.name, + "parkName": _defuse(ident.park.name), "entityId": ref.id, - "entityName": ref.name, + "entityName": _defuse(ref.name), "entityType": ref.entity_type, - "date": row.date.isoformat(), - "firstOperatingAt": row.firstOperatingAt.isoformat() if row.firstOperatingAt else "", - "lastClosedAt": row.lastClosedAt.isoformat() if row.lastClosedAt else "", - "operatingMinutes": row.operatingMinutes, - "downMinutes": row.downMinutes, - "showCount": row.showCount if row.showCount is not None else "", - "changes": row.changes, - "standbyMin": standby.min if standby else "", - "standbyP50": standby.p50 if standby else "", - "standbyMean": standby.mean if standby else "", - "standbyP90": standby.p90 if standby else "", - "standbyMax": standby.max if standby else "", - "singleRiderP50": single.p50 if single else "", - "singleRiderMax": single.max if single else "", + **_cells(row), } @@ -150,26 +309,33 @@ def __init__(self, handle: TextIO, fmt: str, write_header: bool, ident: _RowIden self._handle = handle self._fmt = fmt self._ident = ident - self._csv = None - if fmt == "csv": - self._csv = csv.DictWriter(handle, fieldnames=CSV_COLUMNS) - if write_header: - self._csv.writeheader() + if fmt == "csv" and write_header: + # A UTF-8 BOM, so Excel on Windows does not read the file in the local + # code page and render `Walt Disney World® Resort` as mojibake. The + # primary reader of this file is a spreadsheet. Written with the header, + # so a resumed file never gains a second. + handle.write("\ufeff" + _csv_line(dict(zip(CSV_COLUMNS, CSV_COLUMNS)))) def write(self, ref: EntityRef, row: Any) -> None: - if self._csv is not None: - self._csv.writerow(_csv_row(ref, row, self._ident)) + if self._fmt == "csv": + self._handle.write(_csv_line(_csv_row(ref, row, self._ident))) return # Identity keys come FIRST in the object, so a human reading one line of # NDJSON sees what it is before the numbers. - payload = { + identity = { "parkId": self._ident.park.id, "parkName": self._ident.park.name, "entityId": ref.id, "entityName": ref.name, "entityType": ref.entity_type, - **row.model_dump(mode="json"), } + payload = {**identity, **row.model_dump(mode="json")} + # THE ROW DOES NOT GET TO RENAME THE RUN. Models keep undeclared fields + # now, so a row carrying its own `entityId` or `parkName` would win the + # merge and silently relabel every line of a paid export. Re-asserting + # after the merge costs nothing: a dict keeps the position of the FIRST + # insertion, so identity still reads first. + payload.update(identity) self._handle.write(json.dumps(payload) + "\n") @@ -252,8 +418,44 @@ def _window_floor(exc: APIError) -> str | None: # So state is explicit: the format, the range asked for, the high-water day, and # whether it finished. Unreadable state is treated as no state rather than # crashing -- a corrupt file must not be a permanent wall. +# THE FORMAT IS IN THE FILENAME, not only inside the file. One state file served +# both formats, so finishing a park in ndjson, then csv, then ndjson again left +# the ndjson file with no state describing it -- `resuming` false, `has_rows` +# false, opened in append mode, every row duplicated, exit 0. That is verbatim +# the defect the state file was introduced to prevent. STATE_SUFFIX = ".backfill-state.json" +#: Bumped when the meaning of a field changes. A state file from another version +#: is refused rather than guessed at. +STATE_VERSION = 1 + +SDK_NAME = "py" + + +def state_path_for(out_dir: Path, park_id: str, fmt: str) -> Path: + """Where this park's state lives, for this format.""" + return out_dir / f"{park_id}.{fmt}{STATE_SUFFIX}" + + +def _columns_fingerprint(fmt: str) -> str: + """A short hash of the exact header this build writes. + + THE HEADER IS PART OF THE RESUME CONTRACT and nothing recorded it. 3.3.0 wrote + 19 columns in a different order; 4.0.0 writes 41. `_decide` compared only the + format, so a 3.3.0 state file resumed happily and 41-field rows were appended + under a 19-column header: pandas refuses the file outright, and + `csv.DictReader` silently reads standbyMin as changes. Exit 0 either way. + + Deriving the columns from the schema removed the reviewed diff that used to + make a header change visible, so the fingerprint is what replaces it: any + change to the column list, from any cause, makes every in-flight resume + refuse instead of corrupt. + """ + if fmt != "csv": + return "" + digest = hashlib.sha256("\n".join(CSV_COLUMNS).encode("utf-8")).hexdigest() + return digest[:16] + def _read_state(path: Path) -> dict[str, Any]: try: @@ -267,6 +469,29 @@ def _write_state(path: Path, **fields: Any) -> None: path.write_text(json.dumps(fields, sort_keys=True) + "\n", encoding="utf-8") +def _state_mismatch(state: dict[str, Any], fmt: str) -> str | None: + """Why this state file cannot be resumed by this build, or None. + + camelCase keys, deliberately: the JavaScript SDK writes the same file and the + two used to differ only in `last_day` vs `lastDay` -- the two keys that matter + on the interrupted path. Everything else was spelled identically, so the safe + paths interoperated and nothing warned, while a Python run interrupted at 64 + rows and resumed by the JavaScript command produced 172 rows with 64 + duplicated keys and `complete: true`. + """ + if state.get("stateVersion") != STATE_VERSION: + written = state.get("stateVersion") + return f"it was written by a different version of this command (state v{written})" + if state.get("sdk") != SDK_NAME: + other = state.get("sdk") + return f"it was written by the {other} SDK, and resuming across SDKs is not supported" + if state.get("format") != fmt: + return f"it is a {state.get('format')} run" + if state.get("columns") != _columns_fingerprint(fmt): + return "the column layout changed since it was written" + return None + + class _StateFile(NamedTuple): """Where the state lives and the range it describes, fixed for one park.""" @@ -276,7 +501,9 @@ class _StateFile(NamedTuple): end: Day -def _record(sf: _StateFile, last_day: date | None, *, complete: bool) -> None: +def _record( + sf: _StateFile, last_day: date | None, resume_from: str | None, *, complete: bool +) -> None: """Write the state file. `complete` is the fact the old checkpoint could not express. `sf.start` is the ORIGINAL start of the range, not the day a resumed run @@ -286,14 +513,26 @@ def _record(sf: _StateFile, last_day: date | None, *, complete: bool) -> None: """ _write_state( sf.path, + sdk=SDK_NAME, + sdkVersion=PACKAGE_VERSION, + stateVersion=STATE_VERSION, format=sf.fmt, - start=str(sf.start), - end=str(sf.end), - last_day=last_day.isoformat() if last_day else None, + columns=_columns_fingerprint(sf.fmt), + # None, never the string "None". `str(None)` put the literal "None" in + # the file where the JavaScript SDK writes null, and "None" is truthy. + start=_day_str(sf.start), + end=_day_str(sf.end), + lastDay=last_day.isoformat() if last_day else None, + resumeFrom=resume_from, complete=complete, ) +def _day_str(value: Day) -> str | None: + """A day as a string, or None. Never the string "None".""" + return None if value is None else str(value) + + Day = Union[str, date, None] @@ -319,12 +558,30 @@ def _say_empty(end: Day, start: Day | None = None) -> None: ) +def _next_page_start(next_url: str | None) -> str | None: + """The day a paged history URL starts on, or None if it does not say. + + The SDK follows the server's `next` verbatim; all that is wanted here is its + `from`, to write into the state file as a bare day. A day keeps the state + readable and sends a resumed run down the same code path as a first run. + """ + if not next_url: + return None + query = parse_qs(urlparse(next_url).query) + values = query.get("from") or [] + return values[0] if values and values[0] else None + + class _Plan(NamedTuple): """What a run should do about a park, once its state file has been read.""" start: Day has_rows: bool prior_start: str | None + #: True when rows from an EARLIER run are already in the file. Every deletion + #: in this module has to consult it: `written == 0` means "this process wrote + #: nothing", which on a resumed run is not the same as "the file is empty". + resumed: bool = False def _decide( @@ -343,10 +600,11 @@ def _decide( state_path.unlink(missing_ok=True) file_exists = out_path.exists() and out_path.stat().st_size > 0 - same_format = state.get("format") == fmt + mismatch = _state_mismatch(state, fmt) if state else None + resumable = bool(state) and mismatch is None # Finished already. Say so and stop, rather than appending a second copy. - if state.get("complete") and same_format and file_exists: + if state.get("complete") and resumable and file_exists: print( f" already complete: {state.get('start')} .. {state.get('end')} " f"in {out_path.name} — pass --overwrite to fetch it again", @@ -365,22 +623,37 @@ def _decide( ) return 1 - # A format switch cannot resume: the half-written file is the other format. - if state and not same_format and not state.get("complete"): - other = state.get("format") + # A state file this build cannot resume. Refusing is the only safe answer: the + # file beside it was written to a different contract, and appending to it + # produces a file no reader can parse -- or worse, one that parses wrongly. + if state and mismatch is not None and not state.get("complete"): print( - f" {other} was interrupted part-way for this park. Finish it in " - f"{other}, or pass --overwrite to start again in {fmt}", + f" there is an unfinished {out_path.name} beside this state file, but " + f"{mismatch}.\n" + f" --overwrite start this park again from the beginning\n" + f" or move both files aside and run again", file=sys.stderr, ) return 1 - resuming = bool(state) and same_format and not state.get("complete") - last_written = state.get("last_day") if resuming else None + resuming = resumable and not state.get("complete") + # THE PAGE BOUNDARY, not the newest row. `last_day` is the highest date + # written; the page it came from covered further, because an entity that + # stopped reporting has no rows for the tail days. Resuming at `last_day` + # re-fetches a day already in the file and appends every row of it again -- + # on the exit-75 path, which is the ordinary path for a long back fill, and + # it breaks the (entityId, date) key the file is documented to have. + # + # `last_day` stays as the fallback for the two cases with no boundary + # recorded: a state file written by 3.3.0, and a run that died part-way + # through its FIRST page. One duplicated day beats starting from the top and + # appending a second copy of the whole archive. + resume_at = (state.get("resumeFrom") or state.get("lastDay")) if resuming else None return _Plan( - start=last_written or archive_from, + start=resume_at or archive_from, has_rows=file_exists and resuming, prior_start=state.get("start") if resuming else None, + resumed=resuming, ) @@ -409,6 +682,7 @@ class _Progress: def __init__(self) -> None: self.written = 0 self.last_day: date | None = None + self.resume_from: str | None = None self.skipped = False @@ -420,9 +694,18 @@ def _stream(job: _Job, start: Day, progress: _Progress) -> None: a reader had to. This is the streaming. """ + def note_page(page: HistoryPage) -> None: + """Checkpoint, called once every row of a page is written. + + The day the NEXT page starts on, taken from the server's own `next` URL, + so a resumed run asks for nothing twice. None on the last page, where + there is nothing left to carry on from. + """ + progress.resume_from = _next_page_start(page.next_url) + def write_rows(writer: Writer, first_day: Day) -> None: """Stream one range into the file. Raises whatever the SDK raises.""" - for ref, row in job.history.days_with_entities(first_day, job.end): + for ref, row in job.history.days_with_entities(first_day, job.end, on_page=note_page): writer.write(ref, row) progress.written += 1 # MAX, not last-seen. `_daily_rows` walks entities and then each @@ -465,6 +748,45 @@ def write_rows(writer: Writer, first_day: Day) -> None: write_rows(writer, floor) +def _window_closed(out_path: Path, end: Day, start: Day, *, resumed: bool) -> int: + """This key's window does not reach the day this park would start at. + + NEVER DELETE ROWS AN EARLIER RUN DOWNLOADED. `start` is the resume point on a + rerun, so a key rotated out of a scheduler's environment or a lapsed + subscription used to wipe the partial archive and exit 0 -- the scheduler + logged success -- and then trap: the file gone, the state surviving, + `has_rows` false, and every later run re-entering this branch and exiting 0 + with no data. + """ + _say_empty(end, start) + if resumed: + print( + " the rows already downloaded are left alone. Your plan no longer " + "reaches the day this run would continue from", + file=sys.stderr, + ) + return 1 + out_path.unlink(missing_ok=True) + return 0 + + +def _nothing_written( + out_path: Path, state_path: Path, sf: _StateFile, progress: _Progress, *, resumed: bool +) -> int: + """The park had nothing in this key's window. Tidy up, or refuse to. + + On a first run both files go: an empty file reads as "this park has no + history". On a RESUMED run an earlier run's rows are real and are not ours to + remove, so the state is kept and the exit code says the run did not finish. + """ + if resumed: + _record(sf, progress.last_day, progress.resume_from, complete=False) + return 1 + out_path.unlink(missing_ok=True) + state_path.unlink(missing_ok=True) + return 0 + + def backfill_park( tp: ThemeParks, park: _Park, out_dir: Path, fmt: str, overwrite: bool = False ) -> int: @@ -498,13 +820,13 @@ def backfill_park( ext = "csv" if fmt == "csv" else "ndjson" out_path = out_dir / f"{park_id}.{ext}" - state_path = out_dir / f"{park_id}{STATE_SUFFIX}" + state_path = state_path_for(out_dir, park_id, fmt) end = span.retrievable_through decided = _decide(out_path, state_path, fmt, overwrite, span.archive_from) if isinstance(decided, int): return decided - start, has_rows, prior_start = decided + start, has_rows, prior_start, resumed_run = decided resuming = prior_start is not None sf = _StateFile(state_path, fmt, prior_start or start, end) @@ -514,9 +836,7 @@ def backfill_park( ) if _is_empty_window(start, end): - _say_empty(end, start) - out_path.unlink(missing_ok=True) - return 0 + return _window_closed(out_path, end, start, resumed=resumed_run) ident = _RowIdentity(park) job = _Job(history, out_path, fmt, end, has_rows, ident) @@ -526,22 +846,45 @@ def backfill_park( except BudgetExhaustedError as exc: # The budget is hourly, so a spent one can be most of an hour from # resetting. Record how far we got and exit 75 rather than sleeping. - _record(sf, progress.last_day, complete=False) + _record(sf, progress.last_day, progress.resume_from, complete=False) + if progress.written == 0 and not resumed_run: + # A budget spent before the first page left a 0-byte file that reads + # as "this park has no history". + out_path.unlink(missing_ok=True) wait = exc.retry_after or 0 print( f" budget spent; rerun the same command in {wait / 60:.0f} min to continue", file=sys.stderr, ) return EX_TEMPFAIL + # ORDER MATTERS AND IT BIT ONCE: BudgetExhaustedError subclasses + # RateLimitError, so this clause above the budget one catches it first and + # turns exit 75 into a traceback and exit 1 -- the precise regression the + # budget handler exists to prevent. + except (ThemeParksError, OSError): + # Every other failure still records where it got to, or the next run + # starts over and appends a second partial copy. And AN EMPTY FILE IS A + # LIE: opening the file created it before the first request, so a park + # that failed with nothing written left a 0-byte file that reads as + # "this park has no history" -- on a six-park destination the customer + # counts six files and never sees which one is empty. + if progress.last_day is not None: + _record(sf, progress.last_day, progress.resume_from, complete=False) + # `written` counts rows THIS process wrote, so on a resumed run it is 0 + # while the file holds everything the previous runs fetched. Deleting it + # there destroyed the archive and left the state file pointing into the + # middle of it, so the next run appended only the tail and recorded + # `complete: true`. + if progress.written == 0 and not resumed_run: + out_path.unlink(missing_ok=True) + raise if progress.skipped and progress.written == 0: - out_path.unlink(missing_ok=True) - state_path.unlink(missing_ok=True) - return 0 + return _nothing_written(out_path, state_path, sf, progress, resumed=resumed_run) # Completion is RECORDED, never inferred from a missing file. That is the # distinction the old checkpoint could not make. - _record(sf, progress.last_day, complete=True) + _record(sf, progress.last_day, None, complete=True) print(f" done: {progress.written} rows -> {out_path}", file=sys.stderr) return 0 @@ -638,34 +981,58 @@ def parks_in(did: str) -> list[tuple[str, str]]: if len(exact_park) == 1: return exact_park - park_hits = [(pid, pname) for pid, pname, _, _ in catalogue if lowered in _normalize(pname)] + # WHEN THE EXACT NAME IS AMBIGUOUS, the exact matches ARE the candidates. + # "Disneyland Park" is two live parks, Anaheim and Paris; widening to + # substrings adds Hong Kong Disneyland Park, which is not what was typed and + # pads the one list whose whole job is "which of these did you mean". + park_hits = ( + exact_park + if len(exact_park) > 1 + else [(pid, pname) for pid, pname, _, _ in catalogue if lowered in _normalize(pname)] + ) dest_hits = {did: dname for _, _, did, dname in catalogue if lowered in _normalize(dname)} if len(dest_hits) == 1 and not park_hits: return parks_in(next(iter(dest_hits))) - # THE DESTINATION GOES IN THE LABEL, and it is load-bearing: TWO parks are - # named exactly "Disneyland Park" -- Anaheim and Paris -- so a list of bare - # park names offers a choice between two identical lines. - dest_of = {pid: dname for pid, _, _, dname in catalogue} - candidates = [(pid, f"{pname} ({dest_of[pid]})") for pid, pname in park_hits] or [ - (did, f"{dname} (destination, {len(parks_in(did))} parks)") - for did, dname in dest_hits.items() - ] - if not candidates: + # A UNIQUE SUBSTRING RESOLVES. `themeparks-backfill "magic kingdom"` names + # exactly one park, and refusing a query that is unambiguous is hostile. + if len(park_hits) == 1: + return park_hits + + if not park_hits and not dest_hits: raise SystemExit( f'no park or destination matching "{wanted}".\n' f" themeparks-backfill --list everything\n" f' themeparks-backfill --list disney the ones matching "disney"' ) - if len(candidates) > 1: - lines = "\n".join( - f" {cid} {name}" for cid, name in sorted(candidates, key=lambda c: c[1]) - ) - raise SystemExit( - f'"{wanted}" matches {len(candidates)}. Pass an id, or the destination' - f" name to get all of its parks:\n{lines}" - ) - return candidates + + # AMBIGUOUS: list the ids and stop. THE LABEL IS NOT THE NAME -- this used to + # `return candidates`, whose second element is the formatted display label, so + # a one-hit substring downloaded the right park and wrote + # `"Magic Kingdom Park (Walt Disney World® Resort)"` into the parkName column + # of every one of ~73,000 rows, and echoed the destination twice. Four live + # names reach this path. Shipped in 3.3.0. + # + # The destination is load-bearing in the LABEL: two live parks are named + # exactly "Disneyland Park" -- Anaheim and Paris -- so bare names would offer + # a choice between two identical lines. Sorted by park name, id first, so the + # id is the copy-pasteable part. + dest_of = {pid: dname for pid, _, _, dname in catalogue} + labelled = ( + [(pid, pname, f"{pname} ({dest_of[pid]})") for pid, pname in park_hits] + if park_hits + else [ + (did, dname, f"{dname} (destination, {len(parks_in(did))} parks)") + for did, dname in dest_hits.items() + ] + ) + lines = "\n".join( + f" {cid} {label}" for cid, _, label in sorted(labelled, key=lambda c: (c[1], c[0])) + ) + raise SystemExit( + f'"{wanted}" matches {len(labelled)}. Pass one of these ids, or the ' + f"destination name to get all of its parks:\n{lines}" + ) def _resolve(catalogue: list[tuple[str, str, str, str]], wanted: str) -> list[tuple[str, str]]: @@ -694,25 +1061,29 @@ def _looks_like_id(value: str) -> bool: return len(value) == UUID_LENGTH and value.count("-") == UUID_DASHES -def _print_list(tp: ThemeParks, needle: str | None) -> int: +def _print_list(catalogue: list[tuple[str, str, str, str]], needle: str | None) -> int: """Parks grouped under their destination, so a destination id is visible too.""" - catalogue = _catalogue(tp) + shown = catalogue if needle: lowered = _normalize(needle) - catalogue = [ - c for c in catalogue if lowered in _normalize(c[1]) or lowered in _normalize(c[3]) - ] - if not catalogue: + shown = [c for c in catalogue if lowered in _normalize(c[1]) or lowered in _normalize(c[3])] + if not shown: print(f'nothing matching "{needle}"', file=sys.stderr) return 1 by_dest: dict[tuple[str, str], list[tuple[str, str]]] = {} - for pid, pname, did, dname in catalogue: + for pid, pname, did, dname in shown: by_dest.setdefault((did, dname), []).append((pid, pname)) for (did, dname), parks in sorted(by_dest.items(), key=lambda kv: kv[0][1]): - # The destination line is indented left of its parks and labelled, so it - # reads as "pass this to get all of them" rather than as another park. - print(f"{did} {dname} <- destination: all {len(parks)} parks") + # THE TOTAL, counted from the UNFILTERED catalogue. Counting the filtered + # rows made `--list epcot` print "all 1 parks" for a destination with + # six, on the one line whose entire job is that number -- and that line + # is an instruction to pass the destination id, so the number is what the + # reader decides on. + total = sum(1 for c in catalogue if c[2] == did) + note = "" if total == len(parks) else f" ({len(parks)} shown)" + word = "park" if total == 1 else "parks" + print(f"{did} {dname} <- destination: all {total} {word}{note}") for pid, pname in sorted(parks, key=lambda p: p[1]): print(f" {pid} {pname}") return 0 @@ -789,6 +1160,12 @@ def main(argv: list[str] | None = None) -> int: " that finished is not fetched twice and a file this command did not" " write is never touched.", ) + parser.add_argument( + "--version", + action="version", + version=f"themeparks-backfill {PACKAGE_VERSION}", + help="print the version and exit", + ) parser.add_argument( "--out", type=Path, @@ -804,7 +1181,7 @@ def main(argv: list[str] | None = None) -> int: # key and before you have decided to pay for anything. if args.list_parks is not None: with ThemeParks(api_key=args.api_key, user_agent=USER_AGENT) as tp: - return _print_list(tp, args.list_parks or None) + return _print_list(_catalogue(tp), args.list_parks or None) if not args.parks: parser.error("which park or destination? try: themeparks-backfill --list disney") @@ -871,14 +1248,79 @@ def main(argv: list[str] | None = None) -> int: file=sys.stderr, ) - for park_id, pname in targets: + return _run_all(tp, targets, args) + return 0 + + +def _run_all(tp: ThemeParks, targets: list[tuple[str, str]], args: Any) -> int: + """Back fill every target, and report what did not finish. + + ONE PARK'S FAILURE IS NOT THE DESTINATION'S. A 500 on Animal Kingdom used to + abandon the run, so the parks after it were never attempted: the customer got + a partial download, a traceback, and no statement of which parks were + missing. Every park is tried, what failed is named at the end, and the exit + code still says something went wrong. + """ + failed: list[str] = [] + for park_id, pname in targets: + try: status = backfill_park(tp, _Park(park_id, pname), args.out, args.format, args.overwrite) - if status != 0: - # Stop at the first exhausted budget. Carrying on to the next - # park only spends the retry-after on 429s. - return status + except (ThemeParksError, OSError) as exc: + print(f"{park_id}: {exc}", file=sys.stderr) + failed.append(pname) + continue + if status == EX_TEMPFAIL: + # A spent budget stops everything: the next park would spend the + # retry-after for nothing, and every state file says where it got to. + return status + if status != 0: + failed.append(pname) + if failed: + print( + f"\n{len(failed)} of {len(targets)} did not finish: {', '.join(failed)}\n" + f" the rest are written. Run the same command again to retry just these.", + file=sys.stderr, + ) + return 1 return 0 +def cli() -> int: + """The installed entry point: `main`, with no traceback for a bad request. + + An unreachable API, a timed-out connection or a mistyped id used to print a + nine-frame traceback. A traceback is a bug report about this command; none of + these are bugs in it, and a customer who has just paid reads one as the tool + being broken. + """ + try: + return main() + except KeyboardInterrupt: + print("\nstopped. Run the same command again to continue.", file=sys.stderr) + return 130 + except (NetworkError, ApiTimeoutError) as exc: + # A connection reset or a timeout IS resumable, so this is EX_TEMPFAIL and a + # scheduler retries rather than alerting. The JavaScript SDK said 75 here + # and Python said 1: opposite semantics for one event, on the number a cron + # acts on. + print( + f"{exc}\n the run is resumable: the same command continues it", + file=sys.stderr, + ) + return EX_TEMPFAIL + except ThemeParksError as exc: + # Anything the API actively rejected: a bad range, a missing id, a revoked + # key. Retrying changes nothing, so this is a plain failure. Every SDK + # failure descends from ThemeParksError, so a new kind cannot become a + # traceback. + print(f"{exc}", file=sys.stderr) + return 1 + except OSError as exc: + # A full disk or an unwritable --out directory. Five years of one park is + # a few hundred MB, so this is not hypothetical. + print(f"cannot write the output: {exc}", file=sys.stderr) + return 1 + + if __name__ == "__main__": - raise SystemExit(main()) + raise SystemExit(cli()) diff --git a/uv.lock b/uv.lock index 386cada..bfc6c00 100644 --- a/uv.lock +++ b/uv.lock @@ -2114,7 +2114,7 @@ wheels = [ [[package]] name = "themeparks" -version = "3.2.0" +version = "4.0.0" source = { editable = "." } dependencies = [ { name = "eval-type-backport", marker = "python_full_version < '3.10'" },