Skip to content

Repository files navigation

MBTA Delay Estimator

CI

Live at mbta.frankbs.dev.

Real-time map of Boston transit vehicles. Delays are computed from each vehicle's physical position against the published timetable, not read from the agency feed. Over a weekday morning peak the computed figure lands within a minute of the MBTA's own prediction 90% of the time, with a median difference of nine seconds; see Validation.

Live map of Boston with vehicles colored by delay and moving between polls. Hovering a bus shows its delay, searching for route 39 narrows the map to that route, and the panel compares the position-derived figure to the MBTA's predictions

GTFS-realtime protobuf feeds are polled into PostGIS by a FastAPI service; a React + MapLibre frontend consumes the REST API. The MBTA's realtime feeds require no API key.

Stack: Python 3.13, FastAPI, asyncpg, PostgreSQL 17 + PostGIS 3.6, React 18, MapLibre GL, Vite.

Data: MBTA GTFS and GTFS-realtime feeds, provided by MassDOT under its developer license. This project is not affiliated with the MBTA or MassDOT.

Deriving delay from position

Two fields are missing from the MBTA's data:

  • TripUpdates has no delay field, only absolute predicted arrival times.
  • shape_dist_traveled is empty on every shape point in the feed, so there is no published measure of how far along a route a given stop sits.

The second is the bigger problem: without it there's no way to say "this vehicle is between stops 7 and 8, 40% of the way along". app.offsets derives it instead: ST_LineLocatePoint computes, for every distinct (shape, stop) pair, the fraction along the route line at which that stop falls. That turns the timetable into a mapping from position to time, so a live vehicle projected onto its own shape can be compared with when the schedule expected a vehicle at that point.

The work is keyed on (shape_id, stop_id) rather than (trip_id, stop_sequence). In the Fall 2026 feed, 122,478 trips share only 1,147 shapes, which reduces the geometry operations from 3.2M to roughly 24,000.

All distance computation runs in EPSG:26986 (NAD83 / Massachusetts Mainland, in meters) rather than WGS84 degrees, which would bias placement east-west at Boston's latitude.

Placing a vehicle on its route

ST_LineLocatePoint returns the first nearest point on the line, so on a loop route a vehicle on its second pass resolves to a position near the start. The feed's current_stop_sequence picks which leg the vehicle is on, and the fraction locates it along that leg. Each observation records how it was placed:

Method Description
interpolated In transit; scheduled time prorated along the shape between two stops
stopped_at Stopped at a mid-route stop; the deviation at the moment it arrived, held for the dwell
layover Stopped at the trip's first stop; measured against scheduled departure, floored at zero
first_stop Approaching the first stop, with no preceding stop to interpolate from; floored at zero

A vehicle sitting at a stop is measured once, when it first reports itself stopped, rather than reading one second later for every second it dwells. Observations far from their shape or implausibly late are kept but flagged low confidence and left out of the analytics. docs/design.md has the full placement and confidence rules.

Validation

Since the MBTA publishes no delay field, its figure is derived for comparison from the same trip and stop: predicted arrival against scheduled arrival, or predicted departure against scheduled departure for a vehicle on layover. They answer different questions (ours is how late a vehicle is right now, theirs is how late it will be on arrival), so they diverge most at peak service.

Measured on Wednesday, September 16, 2026, thinned to one observation per vehicle per minute:

Morning peak (07:00–09:10) Overnight (00:17–05:00)
Paired observations 80,535 14,032
Distinct vehicles 831 358
Median divergence +9s +7s
p10 / p90 −12s / +52s −12s / +42s
Within 60s of feed 90.0% 93.0%
Within 120s of feed 96.6% 98.1%

By mode

The fleet figure is mostly buses. Over Monday, September 28, 2026, a full service day thinned the same way:

Subway Bus Light rail Commuter rail
Paired observations 37,650 407,297 2,378 32,296
Distinct vehicles 63 760 5 70
Median divergence +2s +8s +4s +18s
p10 / p90 −7s / +42s −15s / +43s −4s / +53s −18s / +136s
Within 60s of feed 95.0% 91.4% 91.5% 67.9%
Within 60s, between stops only 93.3% 85.5% 90.4% 58.7%

Light rail here is the Mattapan line alone: the Green Line runs as added trips, which since September 30 are measured against the nearest scheduled slot (see Known limitations). Over its first morning it agreed with the feed 90.1% of the time within 60s, and scored against actual arrivals it shows the same shape as the other modes (docs/validation.md). Ferries (13 boats, 49% within 60s) are left out: boats don't follow the drawn line.

Commuter rail disagrees because the MBTA's predictions there run optimistic between stations, and its stations are far apart. Scoring both figures against when trains actually arrived (9,510 commuter rail arrivals over September 24–29, less a two-and-a-half-day gap while scoring was stalled) puts the MBTA's prediction 45s early on average when a train is five minutes out and 82s early at ten; ours is within 11s of unbiased at both. The two figures are compared at the same moment, so that gap is the divergence. Neither is simply better: the MBTA's mean error is lower inside five minutes (54s against 68s), ours beyond ten (86s against 92s), and on stopped trains they agree 99% of the time. On buses the MBTA's mean error is lower at every horizon measured. Ours is the naive forecast, the current delay carried to the next stop; the point is how far the MBTA's countdown drifts from even that. The lines with the closest stations agree best (Needham 82%, Fairmount 80%), the long stretches to Fall River least (55%). The same optimism shows on every mode, but a bus is never more than a minute or two from its next stop, so it stays small there. The horizon-by-horizon tables are in docs/validation.md.

The comparison has caught three estimator bugs so far. Scoring the same observations before and after the two most recent fixes, the dwell hold and its cap, moved the share within 60s at peak from 85.0% to 90.0%, cut the stopped-at class's mean absolute divergence from 30s to 10s, and brought the spread from σ 67s to 63s. docs/validation.md has the before-and-after tables, the cap values tried, the breakdown by service level and placement method, and the earlier first-stop bug.

API

Endpoint Returns
GET /api/vehicles Live positions as GeoJSON, with both delay figures
GET /api/vehicles/{id}/history Breadcrumb trail with per-point delay
GET /api/routes Route list, optionally limited to those currently running
GET /api/routes/{id}/shape Route geometry as GeoJSON
GET /api/analytics/delay-by-route Mean delay per route, computed vs. feed
GET /api/analytics/divergence Agreement statistics, broken down by method
GET /api/analytics/timeline Both series bucketed over time
GET /api/analytics/health Poller liveness and data volume

Interactive documentation at /docs.

How it fits together

flowchart LR
    feeds[("MBTA GTFS-realtime<br/>positions + predictions")] -->|every 15s| poller
    gtfs[("MBTA static GTFS")] -->|weekly| load["load job<br/>gtfs_static, offsets"]
    poller["poller<br/>ingest, estimate, score, prune"] --> db[("PostGIS")]
    load --> db
    browser["React + MapLibre"] --> caddy["Caddy<br/>TLS"] --> nginx["nginx<br/>static files, /api proxy"] --> api["FastAPI<br/>read-only, cached"] --> db
Loading

Only the poller writes realtime rows; the API reads. The two run as separate containers so API replicas never double-poll.

Project layout

backend/app/
  schema.sql        the gtfs_ts() time helper and feed metadata
  schema_static.sql routes, stops, shapes, trips, stop times
  schema_realtime.sql  positions, predictions, observations, arrival scores
  gtfs_static.py    GTFS zip into PostGIS via streaming COPY
  offsets.py        ST_LineLocatePoint stop-position cache
  backfill.py       recompute observations from stored positions
  services/
    realtime.py     GTFS-realtime poller
    delay.py        the schedule join and comparison
    scoring.py      both estimates scored against the arrivals that followed
  routers/          vehicles, routes, analytics
backend/tests/      the estimator run against a synthetic route in PostGIS
frontend/src/
  components/MapView.jsx            MapLibre map, diverging delay scale
  components/DelayByRouteChart.jsx  two-series comparison chart
  lib/delay.js                      color scale shared by map and legend

Running with Docker

docker compose up -d --build           # PostGIS, API, poller, and the frontend on localhost:8080
docker compose run --rm load           # downloads and loads the feed, ~60s; repeat weekly

The API and poller start before the feed has been loaded; vehicles carry no delay until load finishes.

Running locally

Requires PostgreSQL with PostGIS, Python 3.13 (3.11 or later works), and Node 22+.

brew install postgresql@17 postgis python@3.13
brew services start postgresql@17
createdb tracker

cd backend
python3.13 -m venv .venv && .venv/bin/pip install -r requirements.txt
cp .env.example .env
.venv/bin/python -m app.gtfs_static     # downloads and loads the feed, ~60s
.venv/bin/uvicorn app.main:app --port 8010
cd frontend
npm install && npm run dev              # localhost:5173

Tests

cd backend
.venv/bin/pip install -r requirements-dev.txt
.venv/bin/python -m pytest        # estimator rules; creates a tracker_test DB

cd frontend
npm test                          # delay scale, formatting, chart helpers

CI runs both suites and builds the Docker images on every push.

Deployment

The poller writes and the API only reads, so in production they run as separate processes: any number of API replicas with RUN_POLLER=false, and exactly one poller. docs/operations.md covers the commands, the weekly feed reload, and backfilling after an estimator change.

Known limitations

  • The Green Line is measured against a slot, not a trip of its own. The MBTA publishes every Green Line train as an added trip with no timetable, so when one is first seen it is pinned to the scheduled trip on its branch and direction whose timetable is nearest at that stop, and the estimator reads that timetable from then on; the vehicle card says "Nearest scheduled slot" when that is the case. A train that crawls after leaving the terminal reads late correctly; one that missed its slot by a whole headway is matched to the next slot and reads as on time. Replacement shuttles, and added trips with no slot within an hour, are drawn without a figure and the map says so.
  • Loop routes lose some in-transit vehicles. ST_LineLocatePoint resolves a point to its first match along the line, so on a leg whose stops run backwards along the shape (2.6% of trips) a moving vehicle can't be placed and is left unscored rather than scored against the wrong stop. Stopped vehicles on those trips are unaffected.
  • Between stops, the schedule is prorated evenly along the shape. A bus that crawls through the first half of a leg reads late and then recovers by the next stop. That is inherent to the method and the main reason in-transit observations agree with the feed less closely than stopped ones do.
  • A wrong trip assignment in the feed passes through. A vehicle reporting a trip whose schedule is nowhere near it gets a plausible-looking figure that is simply wrong; only results beyond three hours are flagged low confidence. These were the outliers no estimator change could fix in docs/validation.md.
  • The projection is specific to Massachusetts. Targeting another city means changing the SRID, not only the feed URLs.

About

Derives MBTA vehicle delays from GPS position against the timetable using PostGIS, then validates them against the agency's own predictions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages