Pittsburgh Housing Typology, Equity & Climate Matchmaker. AI for Housing Hackathon (AI Horizons 2026, Pittsburgh). Track: Challenge 3: Housing Typology, Equity & Climate Matchmaker.
| Demo video | https://www.youtube.com/watch?v=zWiRUNKKWFk |
| Live app | https://lotline-pgh.vercel.app |
| Repository | https://github.com/leonac24/LotLine |
Decision support only. Lotline is not zoning, legal, or financial advice. Confirm anything consequential with the City of Pittsburgh Department of City Planning.
A community development corporation (CDC) with a vacant lot in mind has to decide which housing concepts are worth paying an architect, a lender and a community process to look at. Today that first screen means reading Title Nine district by district and pulling parcel, hazard, income and cost data from a dozen places. Even then, the options pull against each other: more homes, deeper affordability, lower carbon, less site risk. A tool that names one right answer hides the tradeoff. Lotline shows what each option gives up and how the ranking depends on whose priorities you use.
- CDC project leads in early predevelopment, choosing which concepts to take into design review. This is the primary user.
- City Planning and the URA, weighing options for publicly held land.
- Residents who want to see the tradeoffs, and the value judgments, behind a proposal.
Pick any of the 22,233 vacant lots in the City of Pittsburgh (from county
assessment records). Lotline compares six housing types from
data/config/typologies.yaml (detached house through midrise) by:
- who each one houses and what it costs to build and live in;
- what it emits over 30 years (materials, home electricity and driving);
- what the Zoning Code allows;
- how the ranking changes depending on whose priorities you use.
The workflow:
- Pick a lot. A stylized 3D model of Pittsburgh on real terrain, with every vacant parcel at its real location. Search by address or parcel ID, describe the lot you want in plain English, filter by size, zoning, public ownership or site hazard, or start from suggested lots spread across the city.
- Build on it. Drag housing types onto the lot. The pad uses the lot's real frontage and depth from the deed legal description where available. Buildings that are off the lot, overlapping, or prohibited by zoning turn red. Every change re-evaluates "Your plan" on the server.
- Compare. Your plan next to each housing type: zoning status with Title Nine citations, "Could this household afford it?" for illustrative households, a line-item budget (land from comparable sales, construction, site work, soft costs, financing) for renting and for owning, carbon over time with crossover years, and site flags (slope, landslide, undermining, flood, combined sewer and modeled sewer overflow).
- Whose priorities? A 100-point weight budget and stakeholder presets (long-time resident, CDC, developer, City Planning, URA, climate). SMAA bars show how often each option ranks first once weights and uncertainty are sampled. A one-line "ranking flip" gives the smallest weight change that swaps the top two.
- Work backwards. Pick a housing type, a home count and an income tier. See the zoning rules that fail (with the variance or special exception path), the subsidy gap per home, and the site flags.
- What we don't know. Every estimate, unconnected source and unreviewed zoning district, plus citywide coverage counts.
- Where does this number come from? An (i) next to every score, lot fact and zoning answer opens the method, the live inputs behind that number on this lot (value, range, provenance), its sources and what it doesn't tell you. Zoning answers show the verbatim Title Nine quote.
- Memo. One printable page for a community meeting or a City Planning conversation.
Evidence and values are separate. Every metric carries a range, a provenance and its sources: solid = observed, dashed = modeled, outlined = assumption, hatched = placeholder. Weights are shown in their own color and never enter a metric. An option that fails a verified zoning requirement is excluded from the ranking, not scored low, and is still shown with its citation.
Source: diagrams/lotline-architecture.mmd
(also rendered as SVG, and as an .excalidraw scene you can open at
excalidraw.com).
Three layers, separated on purpose.
Offline pipeline (pipeline/, Python + GeoPandas, never deployed). Adapters
pull WPRDC parcels, assessments, city-owned property, zoning and three hazard
layers; enrich_context.py joins ACS tract income, FEMA flood zones, PWSA
combined sewersheds and EPA Smart Location transit access. The result is
data/processed/parcels.json — 22,183 vacant lots — plus a compact
parcels.geojson for the map. Zoning is a separate track: extract.py sends
saved Title Nine text to Claude, build_rules.py expands 30 reviewed facts into
171 district rules, and the build drops any rule whose quote is not found word
for word in the saved text.
Shared engine (core/). The pipeline, the server and the tests all import
the same code, so there is one definition of what a metric means. config.py
validates twelve YAML files with pydantic at startup and fails loudly.
engine.py computes each metric with a low/high band and a provenance label.
zoning.py resolves a map code into its two halves — the base district decides
what may be built, the subdistrict suffix decides how big — because that is
how Title Nine is written.
Runtime. server/app.py is a FastAPI app served as a single Vercel Python
function through api/index.py. Four modules talk to Claude, and each one
validates the reply before it reaches a user. The browser gets React plus a
three.js lot view, and runs SMAA and the weight sliders locally.
- The map sends the parcel id to
/api/analysis/{id}. core/engine.pybuilds every metric for that lot, each carrying a range, a provenance label and the source ids behind it.core/zoning.pychecks each housing type against the rules in force and marks prohibited uses as excluded, not merely low-scoring.- The response goes to the browser, which computes the ranking, the SMAA robustness bars and the ranking-flip sentence locally.
- Moving a weight slider recomputes step 4 only. It never re-asks the server, because weights are not allowed to touch the evidence.
The decisions we would have to defend, and what each one costs.
Evidence and values are separate layers. Weights live in the browser and
never reach core/engine.py. The cost is real: rankings cannot be precomputed
server-side, and SMAA has to be fast enough to run on every slider drag. The
benefit is that it is structurally impossible for someone's preferences to
contaminate a measurement — not a policy we follow, a path that does not exist.
Provenance floors at the weakest input. weakest() means a metric built
from three observed values and one placeholder is labelled placeholder. This
makes the app look worse than it is: most scored criteria read placeholder
even where real data does most of the work. We kept it because the alternative —
labelling a metric by its best input — is how a tool starts overclaiming.
Config over literals. No Pittsburgh fact, typology, district code, colour or
number is written in code; everything iterates over data/config/. The cost is
indirection, and a test (test_no_hardcoding) that swaps in fake typologies to
prove it. The benefit is that the Pittsburgh content is reviewable in one place,
and extending to another municipality is a data job rather than a rewrite.
AI-extracted zoning rules are in force without a planner's sign-off.
require_human_review: false. 171 rules are live; none has been checked by a
person, so the approval path on 87.9% of vacant lots rests on an unchecked
extraction. We took that trade because the alternative is 0% coverage, and
because every quote is verified verbatim against the saved code text and every
answer is labelled "not checked by a planner". Flipping the flag to true
restores the gate and zeroes the coverage. The honest framing is that this is a
working, tested, human-in-the-loop mechanism that has not yet been used.
Rules are keyed by fact, not by district. Title Nine sets use by base district and size by subdistrict, so 30 facts fan out to 171 rules and roughly 30 extractions cover all 54 zoning map codes. The cost is a two-part lookup and a schema most people would not guess. The benefit is that a use rule is transcribed once rather than five times, and a reviewer checks 30 things instead of 171.
A bad LLM reply is discarded whole, not patched. If any sentence cites an unknown metric id or uses a number absent from the input, the server throws away the entire response and shows a deterministic template. Redacting sentence by sentence would preserve more text, but it puts visibly chewed-up prose on screen and invites the model to smuggle claims past a partial filter.
Proximity is not capacity. We know each lot's combined sewershed and FEMA
flood zone, and we deliberately do not feed either into a scored metric. Being
inside a sewershed is not evidence of spare sewer capacity. They appear as site
facts with that caveat attached, and city.yaml says so in a comment so the
next person does not "fix" it.
Site context is tested at one point. Slope, landslide, undermining, flood and sewershed are point-in-polygon tests against a parcel centroid. Part of a lot can cross a boundary without its tested point doing so. Polygon-overlap shares would be more accurate and much slower to compute citywide.
The parcel index is a static file, not a live query. No database at runtime; the server reads a precomputed JSON index. This makes a parcel click fast and the deployment tiny, at the cost of a rebuild step and a config hash that drifts from the shipped artifacts whenever config changes without one.
A 3D lot view instead of a 2D map. The simulator shows what a housing type would actually put on the lot, which a choropleth cannot. We gave up the zoom, pan and basemap conventions people already know, and every piece of scenery had to be labelled decorative so nobody reads meaning into it.
Working, on real data:
- Citywide vacant-parcel index (22,233 lots) with zoning district, neighborhood, tract, public ownership, and site hazard flags. Every lot also has a per-parcel evidence record: each value with its geography, date, source, provenance and what still has to be confirmed.
- Zoning rules for 7 of 41 base districts, covering 87.9% of vacant parcels. Each rule cites its section and quotes the code verbatim.
- Affordability: HUD FY2026 Pittsburgh income limits; 2024 ACS tract median income, renter cost burden and renter income bins (about 22,100 lots matched). The local affordability gap interpolates the share of nearby renters below each option's required income.
- Finance model: separate rental and ownership scenarios with line-item uses, land from qualified recent vacant-land sales where available, HUD cost limits as references, PHFA contingency and operating guidance, current property and transfer tax rates, Freddie Mac mortgage rates and PWSA fees.
- Carbon for every parcel and housing type: DOE 2021 IECC climate-zone 5A heat-pump prototype electricity, a published A1–A3 materials benchmark, BTS LATCH household driving annualized with NHTS, and the NREL Cambium 2024 grid pathway. It is labelled a partial scenario, not a complete whole-life total.
- Sewer: PWSA combined-sewershed screen plus ALCOSAN Clean Water Plan modeled typical-year outfall overflow (14,948 direct matches, a labelled regional fallback for the rest), shown apart from each option's added wastewater design flow.
- FEMA flood, steep-slope, landslide-prone and undermined-area screens.
- Scoring, SMAA, ranking flip, work backwards, grounded AI explanations, (i) method popovers, memo.
Estimates and gaps (labelled in the app):
- No numerical placeholders remain. The 14 former stand-ins are now sourced models, local donor estimates (lots with no tract, block-group or legal dimension match borrow a median from matched Pittsburgh neighbors), or declared budget scenarios. Hazard cost reserves are budget stress tests, not remediation estimates. Each carries its range and provenance.
- EPA Smart Location transit access (15,553 lots) is a 2021 snapshot, not current route-level commute time. Current PRT GTFS, HUD CHAS and ResStock remain unconnected.
- Overflow data is a 2018 historical model: context, not current overflow or spare sewer capacity.
- Zoning overlays (Riverfront, IPOD, historic) and 34 smaller base districts have no rules. Those lots show "Needs planner review".
The current counts are generated into docs/LIMITATIONS.md
by uv run python -m pipeline.docs.
- Benefits: CDCs get a first screen of a lot in one place, and residents get to see the value judgments behind a proposal rather than just its conclusion.
- Could be harmed by misuse: residents, if a ranking is used to skip community process (it is only the arithmetic result of the chosen weights); owners of privately held lots, if "vacant" is read as "available"; low-income households, if an affordability "Yes" is read as an eligibility or pricing determination.
- We don't claim: that any lot can be built on or any option is permitted; a count of households displaced (we report the share of new homes priced above what nearby renters can pay); that being near a sewer, stop or school means spare capacity; that presets reflect what real organizations want; or that the AI adds facts.
Full statement: docs/LIMITATIONS.md.
- Sources are cited per metric. Every dataset is registered in
data/config/sources.yamlwith publisher, URL, vintage, license and the date we verified it. Every number indata/config/assumptions.yamlcarries a value, a low–high range, a unit, a source and a rationale. - No PII. Owner names and mailing addresses are never fetched. Households are synthetic and labelled "illustrative". Only public data is sent to the Claude API.
- No secrets in the repo.
.envis gitignored;.env.exampleships instead. - Human-in-the-loop path for zoning. Rules are extracted by AI from saved
Title Nine text, and the build refuses any rule whose quote is not found
verbatim in that text. Rules that pass are in force and every answer says
"AI-extracted from Title Nine with a verbatim quote; not checked by a
planner". A planner confirms rules in
pipeline/zoning/REVIEW.md, which drops the label. Settingrequire_human_review: trueindata/config/zoning.yamlmakes only human-confirmed rules count. A persistent banner sends every user to City Planning. - Uncertainty is carried, not hidden. Metrics have ranges, SMAA samples within them, and the ranking is always shown with its robustness.
All are public. Full registry, licenses and verification dates: docs/SOURCES.md.
| Dataset | Where we got it |
|---|---|
| Allegheny County Property Assessments | WPRDC (CKAN API) |
| Parcel Centroids with Geographic Identifiers | WPRDC |
| Allegheny County Property Sale Transactions (qualified vacant-land sales) | WPRDC |
| Allegheny County Parcel Boundaries | Allegheny County GIS (ArcGIS REST) |
| City-Owned Properties | City of Pittsburgh via WPRDC |
| Pittsburgh Zoning Districts | City of Pittsburgh via WPRDC |
| Pittsburgh Code of Ordinances, Title Nine (Zoning Code) | eCode360 (saved text snapshots in data/sources/zoning/) |
| Pittsburgh Neighborhoods | City of Pittsburgh via WPRDC |
| 25% or Greater Slope, Landslide Prone Areas, Undermined Areas | City of Pittsburgh via WPRDC |
| Street Centerlines | City of Pittsburgh via WPRDC |
| PWSA Combined Sewersheds | PWSA via WPRDC |
| ALCOSAN Clean Water Plan Section 4 modeled overflow tables (2018) | ALCOSAN (PDF report) |
| PWSA water and sewer rates | PWSA |
| FEMA National Flood Hazard Layer (Pennsylvania) | FEMA via PASDA |
| HUD FY2026 Income Limits (Pittsburgh HMFA) | HUD User |
| HUD 2024 Total Development Cost limits | HUD |
| 2024 ACS 5-Year Detailed Tables (B19013, B25070, B25118) | U.S. Census Bureau summary files |
| PHFA Core Application and Operating Budget Instructions | PHFA |
| 2026 Pittsburgh property tax rates; city and county realty transfer tax | City of Pittsburgh, Allegheny County |
| Primary Mortgage Market Survey | Freddie Mac |
| EPA Smart Location Database 3.0 transit access | U.S. EPA (ArcGIS REST) |
| EPA passenger-vehicle emissions factor (GHG Emission Factors Hub) | U.S. EPA |
| EPA eGRID2023 (historical reference) | U.S. EPA |
| Residential Prototype Building Models, 2021 IECC Climate Zone 5A | U.S. DOE / PNNL |
| Climate zone by county (PNNL-33270) | PNNL |
| Embodied carbon benchmarks of U.S. single-family homes (2024) | Published study |
| Cambium 2024 annual grid emissions | NREL |
| LATCH 2017 household travel (tract) | U.S. DOT, Bureau of Transportation Statistics |
| 2017 National Household Travel Survey, Table 30 | FHWA |
| Terrain Tiles (for the 3D city only; never in a metric) | AWS Open Data |
Used to build Lotline:
- Claude Code (Anthropic) wrote most of the first-pass code and docs with the team: pipeline, engine, API, frontend and tests. A person on the team decided each design question. Claude Code also hand-extracted the current zoning facts from Title Nine and checked the dataset endpoints.
- GitHub Copilot coding agent authored a few commits exclusively for fixing merge conflicts.
Inside Lotline:
| Where | Model | Guardrail |
|---|---|---|
| Zoning extraction (offline) | Claude API | JSON schema from config; every rule must quote the code verbatim or it is dropped |
| Tradeoff explanations and "Ask about this lot" (runtime) | Claude API | Every sentence must cite input metric IDs and use only input numbers, or the server falls back to a template |
| Plain-English lot search (runtime) | Claude API | Model can only choose from real filter values; code, not the model, finds lots |
| Source passage classification (offline) | Laya (local) | Labels are leads for review only; they cannot change metrics, zoning or rankings |
AI never sets a weight, changes a metric, or ranks options. With
LLM_PROVIDER=none the app runs with no AI at runtime. Details:
docs/AI_DISCLOSURE.md.
- Python runtime (deployed): FastAPI, pydantic, PyYAML, numpy, anthropic (Claude API SDK).
- Python pipeline (local only): pandas, GeoPandas, shapely, pyogrio,
requests, Pillow, Beautiful Soup; Poppler
pdftotextfor the ALCOSAN report. - Python dev: pytest, ruff, uvicorn, httpx. Managed with uv.
- Local document compile: Laya, PyTorch (CPU).
- Web: Vite, React 19, TypeScript (strict), three.js, oxlint; Fredoka and Nunito from Google Fonts.
- APIs and services: Anthropic Claude API, WPRDC CKAN API, Allegheny County ArcGIS REST, AWS Terrain Tiles, Vercel hosting.
All Pittsburgh-specific facts live in data/config/. The code iterates over
whatever that folder declares, so extending to another municipality is a data
job, not a rewrite.
Requirements: uv and Node 20+.
uv sync # Python 3.14 + runtime/dev deps
cp .env.example .env # optional: add ANTHROPIC_API_KEY, set LLM_PROVIDER=anthropic
uv run uvicorn server.app:app --port 8000 --reload
cd web && npm install && npm run dev # http://localhost:5173 (proxies /api to :8000)Tests and lint:
uv run pytest && uv run ruff check .
cd web && npx tsc -b && npm run lintDeploy (Vercel). Import the repo in Vercel. vercel.json builds web/ as
static files and serves api/index.py (FastAPI) as a Python function. Add
LLM_PROVIDER and ANTHROPIC_API_KEY as environment variables.
To rebuild the parcel index from WPRDC:
uv sync --group pipeline
uv run python -m pipeline.build_parcels # ~2 min first run; caches in data/raw/
uv run python -m pipeline.build_terrain # 3D city heightmap -> web/public/data/terrain.*
uv run python -m pipeline.build_basemap # streets, neighborhood names, lot outlines -> web/public/data/basemap/
uv run python -m pipeline.docs # regenerate SOURCES.md + LIMITATIONS.md countsTo re-run official ACS tract estimates and parcel-point FEMA/PWSA site context without rebuilding assessments:
uv sync --group pipeline
uv run python -m pipeline.enrich_context
uv run python -m pipeline.docsTo refresh the carbon and sewer-overflow models from their sources:
uv run python -m pipeline.carbon_prototypes --refresh --download-doe
uv run python -m pipeline.carbon_geography --refresh
uv run python -m pipeline.overflow_context --refresh
uv run python -m pipeline.build_environment --refresh-metadata
uv run python -m pipeline.docsWithout the download flags these rebuild from the checked-in extracts in
data/models/. See docs/ENVIRONMENT_ESTIMATES.md.
The 2024 ACS table-based summary files need no API key. Raw GIS downloads are
cached in data/raw/; remove a source's cached file to fetch its latest version.
Derived facts and coverage counts are checked in. The flood and sewer screens
test a point inside the parcel where a boundary is available, so they can miss
hazards covering only another part of a lot.
Save the Title Nine sections as text in data/raw/zoning/ (eCode360 blocks
scripts), set ANTHROPIC_API_KEY, then run:
uv run python -m pipeline.zoning.extract --districts R1D-H RM-M R2-LRules are in force once their quotes verify. To have a person confirm them,
tick facts in pipeline/zoning/REVIEW.md and run
uv run python -m pipeline.zoning.build_rules --apply.
All Laya tooling, requirements, curated Laya documents, optional local inputs,
and compiled artifacts live under dev/laya/. The app reads the checked-in
dev/laya/compiled/laya_evidence.json index but never installs or invokes Laya.
See dev/laya/README.md for passage classification and
structured candidate extraction. Numeric candidates are source-linked and
unapplied; structured sources such as ACS, parcel geometry, and EPA SLD are
integrated through deterministic adapters. Laya classification does not create
or apply zoning rules.
Quote-verified AI-extracted zoning rules are currently in force while labeled
as not checked by a planner. Use the dedicated Title Nine pipeline and
require_human_review setting to control that behavior.
- Methods: each computation in plain language, with formulas
- Limitations: what it gets wrong and who could be harmed
- Sources: dataset registry, generated from config
- AI disclosure: AI used to build it and inside it
- Backend: API principles and security posture
- Environmental estimates: carbon and overflow models, their scope and how to refresh them
- Evidence pipeline: the per-parcel evidence record, sales index and finance model
- Estimate status: what replaced each former placeholder and what still needs better evidence
- Have a planner check the AI-extracted zoning rules, starting with the
districts with the most vacant lots: R1D, R2, H, R1A, RM. Then turn on
require_human_review. Extend rules to overlays and the remaining districts (LNC, UI, NDI and the Riverfront districts first). - Tighten the estimates that replaced the placeholders: design-specific material quantities and later life-cycle stages, Pittsburgh-weather energy models, a validated 2010-to-2024 tract crosswalk for travel, and project bids, lender terms and operating budgets in place of declared finance scenarios. Evaluate CHAS for income-tier detail.
- EPA Smart Location Database transit access is joined at block-group level (2021 vintage); refresh it with a newer PRT travel-time matrix if current route-level access is needed. Replace the 2018 ALCOSAN overflow model with current overflow data and get written sewer capacity from PWSA and ALCOSAN.
- Make scenarios editable (tenure, affordability mix, parking) and add the advocates + referee explanation mode.
Pilot partners we'd approach: a Pittsburgh CDC (to test the lot-to-memo workflow on a real parcel), the Department of City Planning (to review the zoning rules and own the human-review step), the URA and the Pittsburgh Land Bank (for publicly held lots), Allegheny County (to extend to other municipalities' codes), and PHFA (for affordability and financing assumptions). Because every Pittsburgh fact lives in reviewable YAML, a partner can maintain the data without touching code.
Three people built Lotline during the build window. Who did what below is read
from the commit history (git log --no-merges); everyone also reviewed and
merged each other's work.
Leona Chen (@leonac24): project lead, frontend and 3D city
- Wrote the first working code: config system, citywide vacant-parcel pipeline, evidence engine and API, then the first web app (map picker, compare view, weights and SMAA, work backwards) and the Vercel deploy.
- Built the 3D simulator: the city on real terrain with rivers, bridges, streets and neighborhood names; the drag-and-drop lot sandbox with live plan analysis and reviewed setbacks drawn as the buildable area; the intro screen; tuck-away panels, a phone layout and touch controls.
- The (i) buttons:
data/config/methods.yamlgives every criterion a method, and each score, lot fact and zoning answer opens that method with the live inputs, sources and caveats behind it on the current lot. - Zoning and costs: captured Title Nine text, extracted rules for the R-family, H and P districts, and sourced the construction cost, soft cost and grid emissions assumptions.
- Merged the Claude Code work (see AI tool disclosure) that added AI lot search, "Ask about this lot", grounded tradeoff explanations, the Next steps tab with AI-drafted outreach, share links and the first-run tour, and moved the LLM provider from Gemini to the Anthropic API.
Pranav Singhal (@s9kt): product direction, evidence and data pipeline
- Set the product direction in the first commit: the product-decision record
(.grill/cdc-housing-scenario-comparison.md)
that fixed the CDC as the primary user, and the original project brief
(
CLAUDE.md). - Built the Laya evidence pipeline (
dev/laya/,pipeline/laya_compile.py): it compiles versioned Pittsburgh zoning sources into a checked-in evidence index with a freshness check and tests. A second stage asks the model for numeric candidates that must quote their source and are never applied automatically. - Added
pipeline/enrich_context.py, which joins 2024 ACS tract income and renter burden, frontage and depth estimated from parcel polygons, and EPA Smart Location Database transit access to every vacant lot, replacing placeholders with sourced values (docs/PLACEHOLDER_PIPELINE.md). - Replaced the last 14 placeholders across every parcel: the per-parcel
evidence record and qualified land-sales index
(docs/EVIDENCE_PIPELINE.md); the rental and
ownership finance model (
core/finance_model.py); sourced carbon from DOE prototypes, a materials benchmark, LATCH travel and Cambium grid pathways; and ALCOSAN modeled overflow context (docs/ENVIRONMENT_ESTIMATES.md).
Devin Myers (@Devin-M5706): backend and zoning engine
- Hardened the API: hard requirements before weights, every input limit and
route bound taken from config with tests that keep it that way, the parcel-id
format in
city.yaml, and the auditable contract in docs/BACKEND.md. - Zoning engine: separate use and dimensional rules, a citywide ADU rule, the height rule, and lot fit from building footprints.
- The "what to find out" work plan:
core/inquiries.pyturns each parcel's unknowns into questions, who can answer them and what to ask for, andweb/src/lib/leverage.tsorders them by how much each answer could move the ranking. - Wrote the design for adding tax to the evidence layer.
Built from scratch during the hackathon build window, which opened Saturday, September 26, 2026 at 9:00 a.m. ET. No code predates kickoff, and the commit history is intact. Libraries, datasets and APIs are listed above. The repository stays public after the event.
