Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 16 additions & 21 deletions .github/workflows/spec-sync.yml
Original file line number Diff line number Diff line change
Expand Up @@ -298,15 +298,17 @@ jobs:
# and a separate sync branch, with a V2-tailored wiring prompt. Kept as a distinct job (not a
# matrix) so the proven V1 loop is untouched and the two prompts can diverge freely.
#
# HOSTS (do not conflate): the V2 *spec* is published at aide.[env]/openapi.json; the V2 *API*
# (what the SDK and contract tests call) is api.ade.[env]. We fetch drift from aide; the SDK
# never talks to aide (its paths are staff-SSO-gated).
# SOURCE: the V2 spec is the FULL customer surface aide's publish-openapi.yml generates from
# release/staging and uploads to public S3 — including routes that are not documented yet
# (v2 workflow, v2 build-schema). The gateways' own /openapi.json serves only the documented
# subset, so it is not a drift source. The V2 *API* (what the SDK and contract tests call)
# is api.ade.[env].
spec-sync-v2:
if: github.repository == 'landing-ai/ade-python'
runs-on: ubuntu-latest
timeout-minutes: 60 # headroom for the AI wiring step (up to --max-turns 250) plus rye sync
env:
V2_SPEC_URL: https://aide.staging.landing.ai/openapi.json # staging drives the loop
V2_SPEC_URL: https://ade-specs.s3.amazonaws.com/v2/staging/openapi.json # staging drives the loop
Comment thread
Copilot marked this conversation as resolved.
SYNC_BRANCH: spec-sync/v2 # separate branch so V1 and V2 sync PRs never collide
SPEC_LABEL: V2 # used only in Slack/PR-comment copy to disambiguate the two jobs
steps:
Expand Down Expand Up @@ -339,22 +341,16 @@ jobs:
if: steps.drift.outputs.code == '0'
run: echo "specs in sync; nothing to do."

# An unavailable spec source (check-drift exit 20) is the EXPECTED case on staging, not an
# incident: staging auto-reclaims and must be booked to come back — an unbooked cluster serves a
# 404 (or the host stops answering entirely). Treat ONLY this as a no-op — log and end the run
# cleanly (no drift, no PR, no Slack); the next hourly run picks up any drift once staging is
# booked. Every other failure (a reachable source returning 401/403/5xx, an empty/invalid spec,
# or a script error) falls through to "Fail on spec error" below and alerts.
- name: Spec source unavailable — skip
if: steps.drift.outputs.code == '20'
run: echo "spec source unavailable (exit 20); staging is likely unbooked — skipping this run."
# The V2 source is a published S3 artifact, not a bookable cluster: it does not go away
# when staging is unbooked, so an unavailable source (exit 20) is an incident like any
# other fetch failure, not an expected no-op.

# A reachable-but-invalid spec (empty/whitespace body, malformed JSON) or a script error is a
# real problem — fail loudly so the catch-all alert fires. Distinct from the exit-20 skip above.
# real problem — fail loudly so the catch-all alert fires.
- name: Fail on spec error
if: steps.drift.outputs.code != '0' && steps.drift.outputs.code != '10' && steps.drift.outputs.code != '20'
if: steps.drift.outputs.code != '0' && steps.drift.outputs.code != '10'
run: |
echo "spec fetch/normalize failed (exit ${{ steps.drift.outputs.code }}); staging was reachable but returned an error status (401/403/5xx), an empty/invalid spec, or a script errored."
echo "spec fetch/normalize failed (exit ${{ steps.drift.outputs.code }}); the published S3 spec was unreachable, returned an error status, was empty/invalid, or a script errored."
exit 1

- name: Check for an open sync PR
Expand Down Expand Up @@ -966,9 +962,8 @@ jobs:
text: ${{ steps.ai_commit.outputs.summary }}
thread_ts: ${{ steps.root.outputs.ts }}

# Catch-all failure alert for this job. An unavailable spec source (exit 20) is skipped earlier
# and never reaches here; what this calls out is a reachable source that errored — an HTTP error
# status / empty / invalid spec, or a script error (exit not 0/10/20) — and an AI-wiring crash
# Catch-all failure alert for this job: a spec fetch that failed — the S3 source unreachable,
# an HTTP error status / empty / invalid spec, or a script error (exit not 0/10) — and an AI-wiring crash
# (PR left mechanical-only), with a generic fallback.
- name: Slack — spec-sync failed
if: failure()
Expand All @@ -979,8 +974,8 @@ jobs:
status: failure
title: 'spec-sync ${{ env.SPEC_LABEL }}: run failed'
text: >-
${{ (steps.drift.outputs.code != '' && steps.drift.outputs.code != '0' && steps.drift.outputs.code != '10' && steps.drift.outputs.code != '20')
&& format('Spec fetch/normalize failed (exit {0}) — staging was reachable but returned an error status (401/403/5xx), an empty/invalid spec, or a script errored (an unbooked cluster 404s → exit 20, skipped, not alerted). Investigate.', steps.drift.outputs.code)
${{ (steps.drift.outputs.code != '' && steps.drift.outputs.code != '0' && steps.drift.outputs.code != '10')
&& format('Spec fetch/normalize failed (exit {0}) — the published S3 spec was unreachable, returned an error status, was empty/invalid, or a script errored. Investigate.', steps.drift.outputs.code)
|| (steps.ai_wiring.outcome == 'failure'
&& 'AI wiring step failed. Any open sync PR has only the mechanical commit and needs manual wiring.'
|| 'spec-sync run failed. See the workflow run for details.') }}
Expand Down
14 changes: 8 additions & 6 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,16 +144,18 @@ production spec), is committed but **not wired** into `release.yml` — deferred
"Known gaps" explains what to fix before wiring it.

It runs **two independent loops** (one job each): the **V1** loop tracks the V1 spec against
`specs/v1-ade.json`, and the **V2** loop tracks the V2 spec on the AIDE gateway
(`aide.[env]/openapi.json`) against `specs/v2-aide.json` on a separate `spec-sync/v2` branch. Both
reuse the same scripts. (Note the host split: the V2 *spec* is published at `aide.[env]`, but the
V2 *API* the SDK calls is `api.ade.[env]`.)
`specs/v1-ade.json`, and the **V2** loop tracks the full V2 customer spec aide publishes to
`https://ade-specs.s3.amazonaws.com/v2/staging/openapi.json` against `specs/v2-aide.json` on a
separate `spec-sync/v2` branch. Both reuse the same scripts. (The gateway's own `/openapi.json`
serves only the documented subset, so it is not the V2 drift source; the V2 *API* the SDK calls
is `api.ade.[env]`.)

On each run a loop fetches and normalizes its live spec (`scripts/spec-sync/fetch-normalize.sh`) and
diffs it against its committed snapshot (`scripts/spec-sync/check-drift.sh`). Staging auto-reclaims
diffs it against its committed snapshot (`scripts/spec-sync/check-drift.sh`). For the V1 loop, staging auto-reclaims
and must be booked, so an unavailable spec source (an unbooked cluster 404s, or the host stops
answering) is treated as an expected no-op — the run ends cleanly with no PR and no Slack alert, and
the next run picks up drift once staging is booked. A reachable source that returns an error status
the next run picks up drift once staging is booked. The V2 loop reads a published S3 artifact that
does not depend on staging being booked, so an unavailable V2 source fails and alerts. A reachable source that returns an error status
(401/403/5xx) or an empty/invalid spec still fails loudly and alerts. On drift it opens one PR with
two attributed commits (paths shown for V1; the V2 loop uses the `v2-aide`/`v2_models` equivalents):

Expand Down
3 changes: 2 additions & 1 deletion docs/spec-sync-runbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,8 @@ Silence is also information — but not always: an unavailable staging spec sour
404) is a deliberate **no-op with no Slack alert**, so "no messages" can mean either "no drift" or
"staging is unbooked". A staging cluster that stays unbooked for days silently stops all drift
detection. If the channel has been quiet for an unusually long stretch, check that staging is booked
before assuming the spec is stable.
before assuming the spec is stable. This applies to the V1 loop only: the V2 loop reads a
published S3 artifact, and an unavailable V2 source fails and alerts.

## 2. Reviewing a drift PR

Expand Down
Loading