Skip to content

ci: add a weekly release pipeline with smoke gate and rollback - #629

Merged
omkargaikwad23 merged 36 commits into
mainfrom
feat/release-pipeline
Oct 1, 2026
Merged

omkargaikwad23 merged 36 commits into
mainfrom
feat/release-pipeline

Conversation

@omkargaikwad23

@omkargaikwad23 omkargaikwad23 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

What this does

Adds a weekly release pipeline, .ci/release.cloudbuild.yaml. It tests the image that the Build trigger already pushed for a commit, then deploys it to GKE and Cloud Run. It builds nothing. If no image exists for the commit, the run stops.

Steps

  1. resolve-release: Gets the image digest for the commit and makes the tag vYYYY.MM.DD.<short-sha>. Stops if another release run is active.
  2. release-notes: Lists the PRs merged since the release that GKE serves now. Runs in parallel with the smoke gate. A failure here never fails the release.
  3. smoke-gate: Starts the candidate image and runs three checks over gRPC:
    • Model smoke eval: 6 SQLite prompts (DQL, DML, DDL). The server generates the SQL. verify_release_smoke.py checks coverage, errors, and scores, and retries one time on model failures.
    • Client SQL smoke: client_sql_smoke.py sends fixed SQL with the noop generator and EvalConfig resources, in batch and streaming mode. Golden SQL must score 100, and wrong SQL must score 0.
    • Log scan: Fails on any traceback, ERROR, or CRITICAL line in the server log.
  4. tag-release: Tags the tested digest.
  5. capture-rollback-target: Records the digest that GKE serves now.
  6. deploy-gke: Pins the Deployment to the new digest, confirms that the pods serve it, pings the server over gRPC, and scans the pod log. If a check fails, it pins the previous digest again.
  7. promote-latest: Moves :latest after GKE passes.
  8. copy-cloudrun-image and deploy-cloudrun: Deploy a Cloud Run revision with no traffic, check its digest, shift traffic, and update the jobs.
  9. notify-success: Prints the result line.

Notifications

Each run prints one RELEASE_RESULT line with the status, stage, tag, PR list, and links. alert_policy.yaml defines a log-based alert that sends this line by email. A timeout or a cancel prints no line.

The file has the placeholders TRIGGER_ID and CHANNEL_NAME, so the repo has no project resource IDs. The live policy has the real values.

Design notes

  • eval_client exits 0 for any completed run, so the verifiers grade the CSV reports.
  • Cloud Build has no on-failure hook. For this reason, deploy, check, and rollback share one step, and each failed step prints its own result line.
  • The Deployment has no readiness probe, so the gRPC ping is the health check.
  • The pipeline pins images by digest, so a tag move cannot change the running image.

Testing

In a local container with this branch's .ci/ mounted:

  • Model smoke eval: [PASS] clean run: 6 prompts, 5 scorers
  • Client SQL smoke: [PASS] in batch and in streaming mode
  • Log scan: 0 errors and 0 tracebacks

Setup outside this PR

  • Trigger weekly-evalbench-release (us-central1)
  • Email notification channel and alert policy
  • Change the trigger refs to main after the merge
  • Cloud Scheduler job: Tuesday 08:15 UTC

…ployed service

Deploys the release image to the cluster, waits for the gRPC server, runs a
small NL2SQL evalset end to end, and fails the build if the results regress.

Switches the eval-server deployment to the Recreate strategy so the rollout
does not deadlock on the ReadWriteOnce PVC.
omkargaikwad23 and others added 10 commits September 24, 2026 11:09
Drops the movable promotion tag from tag-release so an aborted run cannot
touch any deployment reference, and adds a step that records the digest the
cluster is currently serving.

The Deployment spec only carries the mutable :latest tag, so the kubelet's
imageID is the only usable rollback target. The step refuses to proceed when
nothing is running or when pods report more than one digest, since neither
state yields a target worth restoring.
Promotes :latest to the smoke-gated digest, restarts the Deployment, and
confirms the pods answer a gRPC Ping before the build passes. Any failure
repoints :latest at the previously captured digest and fails the build.

Deploy, verify, and rollback share one step because Cloud Build has no
on-failure hook -- a separate rollback step would never run.
Copies the digest GKE accepted into the evalbench-dev registry and deploys
it to the Cloud Run service, so both surfaces serve identical bits. The copy
is registry-side: both repos share a host, so gcrane mounts existing blobs.

Deploying by digest doubles as proof the copy landed the right image. Cloud
Run keeps a revision with no traffic until it is ready, so a failed deploy
already leaves the previous revision serving; traffic is pinned back to it
explicitly in case the deploy fails partway.

Renames deploy to deploy-gke now that there are two deploy targets.
The file header still described a pipeline that stopped at tagging, which has
not been true since the deploy steps landed. Trims the remaining comments to
the reason behind each choice and drops the parts that restated the code.
roll_to chains its three commands instead of guarding each with an explicit
return, and the retry loop uses brace expansion rather than spawning seq.
The rollback ignored update-traffic's exit code, so a release that failed to
restore the previous revision reported the same output as one that restored
it cleanly. Names the last known-good revision when the pin fails, matching
how the GKE path reports an unrecoverable rollback.

Adds --quiet so a prompt can never stall the build to its timeout.
gcloud artifacts docker tags add deletes and recreates a tag that already
exists, which needs artifactregistry.tags.delete. The build account holds
artifactregistry.writer, which stops at tags.update. :latest always exists,
so roll_to could never have succeeded.
@omkargaikwad23 omkargaikwad23 changed the title Feat/release pipeline ci: add a weekly release pipeline with smoke gate and rollback Sep 25, 2026
@omkargaikwad23
omkargaikwad23 marked this pull request as ready for review September 25, 2026 07:14
omkargaikwad23 and others added 9 commits September 28, 2026 23:07
* ci(nl2sql): add a QueryData NL2SQL smoke eval and its verifier

Add the configs and the verifier for an NL2SQL presubmit check that runs
QueryData (query_data_api generator) against db_blog on the Cloud SQL
Postgres CI instance, which the skills harness already uses.

- nl2sql_run_config.yaml: 8 DQL prompts, NOOPGenerator, scorers
  returned_sql, executable_sql, set_match and llmrater.
- db_configs/nl2sql_ci_postgres.yaml and model_configs/querydata_ci.yaml:
  all connection values come from env and the password from a secret.
- nl2sql_seed_config.yaml: seeds db_blog once with setup_databases.py. CI
  does not re-seed per run, because a re-seed drops the public schema and
  concurrent builds would break each other.
- verify_nl2sql.py: requires a clean run. All report files, one row per
  prompt, no generator or golden SQL error, SQL on every row, a numeric
  score from every scorer, at least one executed query, and
  pipeline_debug_info in the report. Accuracy is printed but not gated.

cloudbuild.yaml and verifier/verify.py are unchanged. A separate Cloud
Build config will run this check without an Artifact Registry push.

* ci(nl2sql): run the NL2SQL smoke eval as the verify-nl2sql check

Add .ci/nl2sql.cloudbuild.yaml for a verify-nl2sql Cloud Build trigger in
evalbench-testing. Steps: preflight, pull-cache, build-image, nl2sql-eval
and verify.

- The image stays local to the build. Nothing goes to Artifact Registry,
  so PR builds need no registry write access.
- build-image reads the evalbench-harness-ci layer cache but never pushes
  it.
- preflight fails in seconds if a trigger substitution or the database
  secret is empty.
- cloudbuild.yaml and the test/build triggers in cloud-db-nl2sql are
  unchanged.

* fix(query_data_api): make REST token refresh thread-safe

Runner threads share one QueryDataAPIGenerator. When several threads made
the first REST call at the same time, the ones that did not finish refresh()
sent "Authorization: Bearer None" and QueryData returned 401.

A lock now serializes credential creation and refresh. The token is read
while the lock is held. _get_credentials() is renamed to _get_access_token()
because it now returns the token.

The new test sends 8 concurrent first calls. All 8 must send the refreshed
token, and refresh() must run once.
- Gate the image that the Build trigger pushed for the commit with a
  SQLite smoke eval over gRPC. Retry once on model failures.
- Tag the tested digest, deploy it to GKE with automatic rollback, move
  :latest, and deploy it to Cloud Run.
- Print one RELEASE_RESULT line per run and add a log-based email alert.
- Rename verify_nl2sql_smoke.py to verify_release_smoke.py.
Most production clients generate SQL themselves and send it back. The
server uses the noop generator and reads its configs from EvalConfig
resources. The smoke gate tested only server-side generation.

- Add client_sql_smoke.py. It sends fixed SQL through EvalConfig
  resources, the noop generator, and batch and streaming Eval. Golden
  SQL must score 100. Wrong SQL must score 0 on exact_match and
  set_match. No model is called, so a failure blocks the release with
  no retry.
- Fail the gate on ERROR and CRITICAL server log lines, as well as on
  tracebacks.
- Add the exact_match scorer to the model smoke eval.
Eval payloads can quote a traceback. The live GKE pod log had 114 such
quoted tracebacks and 0 real ones. The old scan matched them, so one
client request could roll back a healthy release.

- Add log_problems to lib.sh. It matches ERROR and CRITICAL log records
  and tracebacks only at the start of a line.
- Use it in smoke-gate and deploy-gke. deploy-gke now also fails on
  ERROR records, not only on tracebacks.
- Clarify the release_result comment.
omkargaikwad23 and others added 4 commits September 30, 2026 10:19
Cloud Build rejects $TRIGGER_ID as an unknown substitution, so the
trigger did not start. Match other active runs on the built-in
TRIGGER_NAME, which triggered builds store in their substitutions.

Also list the log scan as the third smoke-gate check, and remove an
overstated claim from the alert policy comment.
# 2. client_sql_smoke.py: the client-generated SQL path, with fixed scores.
# 3. The log scan: any ERROR, CRITICAL, or traceback in the server log
# fails the gate.
- id: smoke-gate

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add a step timeout: on smoke-gate so deploy-gke always has its full budget. Please also add an alert on build status TIMEOUT/CANCELLED/FAILURE

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done: added timeout: '1200s' to smoke-gate, and a build status alert for TIMEOUT, CANCELLED and INTERNAL_ERROR. FAILURE is already covered by the RELEASE_RESULT email, so each run sends one email.

Comment thread .ci/release.cloudbuild.yaml Outdated

echo "===== scan the new pod logs ====="
PROBLEMS=$$(kubectl logs -n "$$NAMESPACE" "$$POD" -c "$$CONTAINER" \
| log_problems)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's no readiness probe, so this pod is already taking production traffic. Any logging.error caused by a client (for example Cannot run set_up_script in eval_service.py) will roll back a healthy release. Could we scan only a startup window or known startup failures, or report without rolling back? A gRPC/exec readiness probe using wait_for_grpc.py would be the real fix; it could be a follow-up.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. The scan now stops at Server started message, so client errors after startup don't roll back. I'll add a readiness probe in a follow-up.

Add a 1200 s timeout to smoke-gate, so a hung eval cannot use the build
time that deploy-gke needs for a rollout and a rollback.

Add build_status_policy.yaml. It emails when a run ends as CANCELLED,
TIMEOUT, or INTERNAL_ERROR, which print no RELEASE_RESULT line. It
matches the build-end audit entry and skips FAILURE, because a failed
step already prints a RELEASE_RESULT line. Each run sends one email.

Put the detail right after the status in the result email. Name the tag
and digest in the success detail, use notify-success as the stage, and
print 'unavailable' when there is no compare link.
The Deployment has no readiness probe, so the new pod takes traffic as
soon as it starts. A client request can then log an ERROR, for example
a missing set_up_script, and roll back a healthy release. Stop the scan
at the "Server started" line. If the line is missing, scan the full log.
@prernakakkar-google

Copy link
Copy Markdown
Collaborator

Note for future PR hygiene: While cohesive as a self-contained .ci/release/ addition, splitting future changes of this size into (1) smoke verification scripts + unit tests and (2) Cloud Build pipeline + monitoring policies makes review and bisection much easier.

Comment thread .ci/release.cloudbuild.yaml Outdated
AND operation.last=true
AND protoPayload.status.code>0
AND protoPayload.status.code!=9
labelExtractors:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Step timeouts produce status.code=9 (FAILURE), not 4 (TIMEOUT): When an individual step hits its step timeout: (for example smoke-gate hitting 1200s), Cloud Build kills the step container immediately (so RELEASE_RESULT is never printed and alert_policy.yaml does not fire) and marks the overall build as FAILURE (protoPayload.status.code=9). Because this filter has protoPayload.status.code!=9, a step timeout won't trigger either alert policy. We should either wrap the smoke-gate body in an inner timeout 1140s ... || fail_release smoke_gate "smoke gate timed out" so RELEASE_RESULT is always logged, or drop protoPayload.status.code!=9 here.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, I removed code!=9, so every failed run now sends an email, including step timeouts.

@omkargaikwad23
omkargaikwad23 merged commit a24bc3b into main Oct 1, 2026
12 of 14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants