Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
133 changes: 133 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,138 @@
# Changelog

## [4.0.0] - 2026-09-28

**3.3.0 was yanked: incomplete CSV export and a resume defect.**

A major version because **the CSV header changed**: fifteen columns were added and
the order is now the schema's, so a reader that takes columns by position gets the
wrong ones rather than an error. Read by name. The library API is backward
compatible.

Everything here came out of porting `themeparks-backfill` to the JavaScript SDK
and then diffing the two outputs over the same park, and out of six reviews of the
result. Two independent implementations reading one API disagree in exactly the
places one of them is wrong. Magic Kingdom's full archive now comes back
**byte for byte identical** from both SDKs: 94,223 rows, 41 columns, the only
differences being today's row, which grows as the day elapses.

### Fixed

- **`themeparks-backfill "magic kingdom"` wrote the wrong park name into every
row.** A name that matched one park by substring returned the formatted display
label, so `parkName` read `Magic Kingdom Park (Walt Disney World® Resort)` for
all ~94,000 rows, and the resolution echo printed the destination twice. Four
live names reached it.

- **The CSV was missing ten of the thirty-six fields the API sends, on every row.**
`unknownMinutes`, the whole `inParkHours` block (the day's numbers limited to the
park's published hours -- usually the ones you want, since a ride "down" at 2am
is not down), `extremeWaits` (how many readings of 480+ minutes are folded into
the statistics, which is how you spot a feed error), and three of `singleRider`'s
five percentiles while `standby` carried all five. On a five-year Magic Kingdom
export, 72,200 of 94,223 rows were missing their in-park statistics. **The column
list is now derived from the model**, so it cannot drift again.

- **Vendored models were stale, and pydantic drops what it does not declare**, so
those three fields were deleted at parse time for every caller of `days()`, not
just for the CSV. Models regenerated, and every model now keeps fields the schema
does not declare (`themeparks._models_base.ApiModel`, `extra="allow"`), so a
field the API adds tomorrow survives parsing and reaches `model_dump()` and the
NDJSON output before this SDK knows it exists. It does not reach the CSV, whose
columns come from the schema.

- **A resumed download duplicated a day.** The checkpoint was the newest row
written; the page it came from covered further, because an entity that stopped
reporting has no rows for the tail days. A rerun re-fetched a day already in the
file and appended every row of it again, breaking the `(entityId, date)` key --
on the exit-75 path, which is the ordinary path for a long back fill. The
checkpoint is now the day the server's own `next` URL starts on.

- **A failure on a resumed run deleted everything already downloaded.** `written
== 0` means "this process wrote nothing", not "the file is empty". The state file
survived pointing mid-archive, so the next run appended only the tail and
recorded `complete: true`. Same for a window that closes under a resumed run --
a key rotated out of a scheduler's environment, a lapsed subscription -- which
additionally exited 0, so the scheduler logged success, and became a permanent
trap.

- **Resuming across versions, formats or SDKs corrupted the file.** One state file
served both formats, so `ndjson` then `csv` then `ndjson` doubled every row in
the first file; and the state carried nothing about the header, so 3.3.0's
19-column file resumed under this build appended 41-field rows beneath it. The
state file is now `<parkId>.<format>.backfill-state.json` and records the SDK,
its version, a state version and a fingerprint of the exact header. Anything that
does not match is refused with a message saying why, never resumed.

- **A network failure or timeout now exits 75, not 1**, so a scheduler retries
rather than alerting; anything the API actively rejected still exits 1. The
JavaScript SDK had these the other way round.

- **A carriage return in an entity name was written unquoted on Python 3.9 and
3.10**, so one row parsed as two with every later column shifted. The `csv`
module's QUOTE_MINIMAL only quotes characters that appear in the line terminator,
and this command sets LF; 3.11 changed the module to always quote CR and LF, so
the defect was invisible on a modern interpreter and live on two supported ones.
The CSV writer now does its own minimal quoting, which also makes the output
byte-identical across Python versions rather than only within one.

- **UTC timestamps are written `Z`, not `+00:00`**, and CSV line endings are LF.
Between them these accounted for 39,201 differing lines against the JavaScript
SDK's output for no difference in meaning.

- **One park's failure no longer abandons the rest of a destination.** Every park
is tried, what failed is named at the end, and the exit code still says something
went wrong. A spent budget still stops everything, deliberately.

- **A failed park no longer leaves a 0-byte file** that reads as "this park has no
history", including when the budget runs out before the first page.

- **A network failure, a full disk or Ctrl-C is a sentence, not a traceback.**

- **The user agent named neither version.** It was the literal
`themeparks-backfill/1`, and it replaced the SDK's own, so a support question had
no version to work from at either end.

- **`--list <text>` reported the wrong total**, printing "all 1 parks" for a
destination with six -- on the one line whose whole job is that number.

- **An ambiguous name listed the wrong candidates**, widening to substrings and
offering a third park that was not what was typed. It now lists the ids of the
parks that actually match, sorted by name.

- **A collection of nested models would have produced phantom columns** and then an
`AttributeError` on the first row. Duplicate column names are now impossible at
import rather than a wrong number under a right-looking header.

- **The NDJSON identity columns could be overwritten by the row** once models kept
undeclared fields.

### Added

- **The CSV carries a UTF-8 BOM**, so Excel on Windows stops rendering
`Walt Disney World® Resort` as mojibake.
- **A cell a spreadsheet would execute is prefixed with an apostrophe** (`=`, `+`,
`-`, `@`, tab, CR). Numeric cells are left alone, so a negative number stays a
number.
- **`on_page` on `days()` and `days_with_entities()`**, called once every row of a
page has been yielded, with a `HistoryPage` (`start`, `end`, `next_url`). The page
boundary is the server's own answer to "where do I carry on", and the rows cannot
tell you.
- **`--version`.**
- `EntityRef` and `HistoryPage` are exported from the package.
- `tests/fixtures/csv_contract.json`, an identical copy of which lives in the
JavaScript SDK. Both suites assert their column list against it, because this is
one command with two implementations and a customer using both should get one
file format.

### Changed

- The `themeparks-backfill` entry point is `themeparks.backfill:cli`, which adds
the top-level error handling. `main()` is unchanged for anyone calling it.
- Model equality and `model_json_schema()` reflect `extra="allow"`: two responses
differing only in an undeclared field now compare unequal, and dumps may contain
fields the schema does not list.

## [3.3.0] - 2026-09-28

### Added
Expand Down
4 changes: 2 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "hatchling.build"

[project]
name = "themeparks"
version = "3.3.0"
version = "4.0.0"
description = "Official SDK for the ThemeParks.wiki API"
readme = "README.md"
requires-python = ">=3.9"
Expand Down Expand Up @@ -34,7 +34,7 @@ dependencies = [
# The archive backfill, as a command rather than a file to copy off GitHub.
# `pip install themeparks` then `themeparks-backfill "Disneyland Park"` is the
# whole path from nothing to a file of history.
themeparks-backfill = "themeparks.backfill:main"
themeparks-backfill = "themeparks.backfill:cli"

[project.urls]
Homepage = "https://api.themeparks.wiki"
Expand Down
26 changes: 25 additions & 1 deletion scripts/regenerate.py
Original file line number Diff line number Diff line change
Expand Up @@ -136,7 +136,12 @@ def apply_nullable_patches(text: str) -> str:
# result was a model that cannot parse the last page of any paged
# response, where `next` is null.
pattern = re.compile(
rf"(?P<head>class {class_name}\(BaseModel\):\n"
# The base class is `ApiModel`, ours, not `BaseModel` -- see
# themeparks/_models_base.py for why. Both are accepted so the patch
# does not silently stop matching if that changes again; the
# invariant check below is what turns a miss into a hard stop, and it
# did exactly that when the base class moved.
rf"(?P<head>class {class_name}\((?:BaseModel|ApiModel)\):\n"
rf"(?:(?: [^\n]*)?\n)*?"
rf" {field_name}:\s*)"
rf"(?P<line>[^\n]+)"
Expand Down Expand Up @@ -182,6 +187,12 @@ def main() -> None:
str(OUTPUT),
"--output-model-type",
"pydantic_v2.BaseModel",
# EVERY MODEL KEEPS WHAT THE SPEC DOES NOT DECLARE. Pydantic's default is
# to drop it, and this spec trails the API by days at a time, so during
# that window new fields were being deleted at parse time for every
# caller -- silently, with nothing failing. See themeparks/_models_base.py.
"--base-class",
"themeparks._models_base.ApiModel",
"--use-schema-description",
"--use-field-description",
"--use-annotated",
Expand Down Expand Up @@ -216,6 +227,19 @@ def main() -> None:
# Auto-format the generated file so `ruff format --check` in CI doesn't
# fail on quote-style or whitespace differences from datamodel-codegen.
print("Formatting generated models with ruff...")
# `check --fix` first, for the import ordering: the custom base class means
# the generator emits a first-party import in among the third-party ones, and
# `format` alone does not sort imports. Without this, `ruff check` in CI fails
# on a file nobody is allowed to edit by hand.
subprocess.run(
# --select I ONLY. Unpinned, this is one `pyproject.toml` edit away from
# deleting imports: the F401 per-file-ignore for `_generated/*` is the only
# reason a stray `import os` survives it today, and this step runs with
# check=False inside the nightly drift job, so a tightened ignore would
# start removing code silently.
[sys.executable, "-m", "ruff", "check", "--select", "I", "--fix", "--quiet", str(OUTPUT)],
check=False,
)
subprocess.run(
[sys.executable, "-m", "ruff", "format", str(OUTPUT)],
check=True,
Expand Down
13 changes: 13 additions & 0 deletions tests/fixtures/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,3 +35,16 @@ It is trimmed to keep, deliberately, every case that broke name resolution:

**Do not edit these by hand.** Re-capture them. If a name upstream has drifted,
that is a real change and the test should notice.

## mk_park_daily_page1.json / mk_park_daily_page2.json

Two consecutive pages of one real request, captured 2026-09-28:
`GET /entity/75ea578a-adc8-4116-a54d-dccb60765ef9/history/daily?from=2026-08-01&to=2026-09-20`
then its `next` followed verbatim. Trimmed to three entities (an attraction, a
show, a restaurant); `range`, `next` and every row are the server's.

They are the oracle for resumable paging. Page one covers through 2026-08-31 and
the server says carry on at 2026-09-01, but two of its three entities have no
rows after 2026-08-30 -- so a checkpoint taken from the newest ROW rewinds and
re-downloads days already written. The same files are in the JavaScript SDK, so
both ports are tested against identical bytes.
58 changes: 58 additions & 0 deletions tests/fixtures/csv_contract.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
{
"_comment": [
"THE CSV CONTRACT, shared by the Python and JavaScript SDKs.",
"Both repos hold an identical copy of this file and assert their own column",
"list against it, because `themeparks-backfill` is one command with two",
"implementations and a customer using both must get one file format.",
"Before this existed, one SDK wrote 32 columns and the other 41, with",
"inParkScheduledMinutes against inParkHoursScheduledMinutes, and the test that",
"claimed to check it was four spot-checks.",
"Generated from the OpenAPI spec: run `npm run regenerate` in the JavaScript",
"SDK, copy this file to both repos, and expect both suites to fail until they",
"agree."
],
"fingerprint": "1b6ea478049dc7ca",
"columns": [
"parkId",
"parkName",
"entityId",
"entityName",
"entityType",
"date",
"firstOperatingAt",
"lastClosedAt",
"operatingMinutes",
"downMinutes",
"unknownMinutes",
"standbyMin",
"standbyP50",
"standbyMean",
"standbyP90",
"standbyMax",
"singleRiderMin",
"singleRiderP50",
"singleRiderMean",
"singleRiderP90",
"singleRiderMax",
"extremeWaitsStandby",
"extremeWaitsSingleRider",
"showCount",
"inParkHoursScheduledMinutes",
"inParkHoursOperatingMinutes",
"inParkHoursDownMinutes",
"inParkHoursUnknownMinutes",
"inParkHoursStandbyMin",
"inParkHoursStandbyP50",
"inParkHoursStandbyMean",
"inParkHoursStandbyP90",
"inParkHoursStandbyMax",
"inParkHoursSingleRiderMin",
"inParkHoursSingleRiderP50",
"inParkHoursSingleRiderMean",
"inParkHoursSingleRiderP90",
"inParkHoursSingleRiderMax",
"inParkHoursExtremeWaitsStandby",
"inParkHoursExtremeWaitsSingleRider",
"changes"
]
}
Loading
Loading