ci: add a weekly release pipeline with smoke gate and rollback - #629
Conversation
…ployed service Deploys the release image to the cluster, waits for the gRPC server, runs a small NL2SQL evalset end to end, and fails the build if the results regress. Switches the eval-server deployment to the Recreate strategy so the rollout does not deadlock on the ReadWriteOnce PVC.
Drops the movable promotion tag from tag-release so an aborted run cannot touch any deployment reference, and adds a step that records the digest the cluster is currently serving. The Deployment spec only carries the mutable :latest tag, so the kubelet's imageID is the only usable rollback target. The step refuses to proceed when nothing is running or when pods report more than one digest, since neither state yields a target worth restoring.
Promotes :latest to the smoke-gated digest, restarts the Deployment, and confirms the pods answer a gRPC Ping before the build passes. Any failure repoints :latest at the previously captured digest and fails the build. Deploy, verify, and rollback share one step because Cloud Build has no on-failure hook -- a separate rollback step would never run.
Copies the digest GKE accepted into the evalbench-dev registry and deploys it to the Cloud Run service, so both surfaces serve identical bits. The copy is registry-side: both repos share a host, so gcrane mounts existing blobs. Deploying by digest doubles as proof the copy landed the right image. Cloud Run keeps a revision with no traffic until it is ready, so a failed deploy already leaves the previous revision serving; traffic is pinned back to it explicitly in case the deploy fails partway. Renames deploy to deploy-gke now that there are two deploy targets.
The file header still described a pipeline that stopped at tagging, which has not been true since the deploy steps landed. Trims the remaining comments to the reason behind each choice and drops the parts that restated the code.
roll_to chains its three commands instead of guarding each with an explicit return, and the retry loop uses brace expansion rather than spawning seq.
The rollback ignored update-traffic's exit code, so a release that failed to restore the previous revision reported the same output as one that restored it cleanly. Names the last known-good revision when the pin fails, matching how the GKE path reports an unrecoverable rollback. Adds --quiet so a prompt can never stall the build to its timeout.
gcloud artifacts docker tags add deletes and recreates a tag that already exists, which needs artifactregistry.tags.delete. The build account holds artifactregistry.writer, which stops at tags.update. :latest always exists, so roll_to could never have succeeded.
…verification script
* ci(nl2sql): add a QueryData NL2SQL smoke eval and its verifier Add the configs and the verifier for an NL2SQL presubmit check that runs QueryData (query_data_api generator) against db_blog on the Cloud SQL Postgres CI instance, which the skills harness already uses. - nl2sql_run_config.yaml: 8 DQL prompts, NOOPGenerator, scorers returned_sql, executable_sql, set_match and llmrater. - db_configs/nl2sql_ci_postgres.yaml and model_configs/querydata_ci.yaml: all connection values come from env and the password from a secret. - nl2sql_seed_config.yaml: seeds db_blog once with setup_databases.py. CI does not re-seed per run, because a re-seed drops the public schema and concurrent builds would break each other. - verify_nl2sql.py: requires a clean run. All report files, one row per prompt, no generator or golden SQL error, SQL on every row, a numeric score from every scorer, at least one executed query, and pipeline_debug_info in the report. Accuracy is printed but not gated. cloudbuild.yaml and verifier/verify.py are unchanged. A separate Cloud Build config will run this check without an Artifact Registry push. * ci(nl2sql): run the NL2SQL smoke eval as the verify-nl2sql check Add .ci/nl2sql.cloudbuild.yaml for a verify-nl2sql Cloud Build trigger in evalbench-testing. Steps: preflight, pull-cache, build-image, nl2sql-eval and verify. - The image stays local to the build. Nothing goes to Artifact Registry, so PR builds need no registry write access. - build-image reads the evalbench-harness-ci layer cache but never pushes it. - preflight fails in seconds if a trigger substitution or the database secret is empty. - cloudbuild.yaml and the test/build triggers in cloud-db-nl2sql are unchanged. * fix(query_data_api): make REST token refresh thread-safe Runner threads share one QueryDataAPIGenerator. When several threads made the first REST call at the same time, the ones that did not finish refresh() sent "Authorization: Bearer None" and QueryData returned 401. A lock now serializes credential creation and refresh. The token is read while the lock is held. _get_credentials() is renamed to _get_access_token() because it now returns the token. The new test sends 8 concurrent first calls. All 8 must send the refreshed token, and refresh() must run once.
- Gate the image that the Build trigger pushed for the commit with a SQLite smoke eval over gRPC. Retry once on model failures. - Tag the tested digest, deploy it to GKE with automatic rollback, move :latest, and deploy it to Cloud Run. - Print one RELEASE_RESULT line per run and add a log-based email alert. - Rename verify_nl2sql_smoke.py to verify_release_smoke.py.
…to include PR comparisons
Most production clients generate SQL themselves and send it back. The server uses the noop generator and reads its configs from EvalConfig resources. The smoke gate tested only server-side generation. - Add client_sql_smoke.py. It sends fixed SQL through EvalConfig resources, the noop generator, and batch and streaming Eval. Golden SQL must score 100. Wrong SQL must score 0 on exact_match and set_match. No model is called, so a failure blocks the release with no retry. - Fail the gate on ERROR and CRITICAL server log lines, as well as on tracebacks. - Add the exact_match scorer to the model smoke eval.
Eval payloads can quote a traceback. The live GKE pod log had 114 such quoted tracebacks and 0 real ones. The old scan matched them, so one client request could roll back a healthy release. - Add log_problems to lib.sh. It matches ERROR and CRITICAL log records and tracebacks only at the start of a line. - Use it in smoke-gate and deploy-gke. deploy-gke now also fails on ERROR records, not only on tracebacks. - Clarify the release_result comment.
Cloud Build rejects $TRIGGER_ID as an unknown substitution, so the trigger did not start. Match other active runs on the built-in TRIGGER_NAME, which triggered builds store in their substitutions. Also list the log scan as the third smoke-gate check, and remove an overstated claim from the alert policy comment.
| # 2. client_sql_smoke.py: the client-generated SQL path, with fixed scores. | ||
| # 3. The log scan: any ERROR, CRITICAL, or traceback in the server log | ||
| # fails the gate. | ||
| - id: smoke-gate |
There was a problem hiding this comment.
Please add a step timeout: on smoke-gate so deploy-gke always has its full budget. Please also add an alert on build status TIMEOUT/CANCELLED/FAILURE
There was a problem hiding this comment.
Done: added timeout: '1200s' to smoke-gate, and a build status alert for TIMEOUT, CANCELLED and INTERNAL_ERROR. FAILURE is already covered by the RELEASE_RESULT email, so each run sends one email.
|
|
||
| echo "===== scan the new pod logs =====" | ||
| PROBLEMS=$$(kubectl logs -n "$$NAMESPACE" "$$POD" -c "$$CONTAINER" \ | ||
| | log_problems) |
There was a problem hiding this comment.
There's no readiness probe, so this pod is already taking production traffic. Any logging.error caused by a client (for example Cannot run set_up_script in eval_service.py) will roll back a healthy release. Could we scan only a startup window or known startup failures, or report without rolling back? A gRPC/exec readiness probe using wait_for_grpc.py would be the real fix; it could be a follow-up.
There was a problem hiding this comment.
Agreed. The scan now stops at Server started message, so client errors after startup don't roll back. I'll add a readiness probe in a follow-up.
Add a 1200 s timeout to smoke-gate, so a hung eval cannot use the build time that deploy-gke needs for a rollout and a rollback. Add build_status_policy.yaml. It emails when a run ends as CANCELLED, TIMEOUT, or INTERNAL_ERROR, which print no RELEASE_RESULT line. It matches the build-end audit entry and skips FAILURE, because a failed step already prints a RELEASE_RESULT line. Each run sends one email. Put the detail right after the status in the result email. Name the tag and digest in the success detail, use notify-success as the stage, and print 'unavailable' when there is no compare link.
The Deployment has no readiness probe, so the new pod takes traffic as soon as it starts. A client request can then log an ERROR, for example a missing set_up_script, and roll back a healthy release. Stop the scan at the "Server started" line. If the line is missing, scan the full log.
…r release pipeline efficiency
…ty and actionable steps
|
Note for future PR hygiene: While cohesive as a self-contained .ci/release/ addition, splitting future changes of this size into (1) smoke verification scripts + unit tests and (2) Cloud Build pipeline + monitoring policies makes review and bisection much easier. |
| AND operation.last=true | ||
| AND protoPayload.status.code>0 | ||
| AND protoPayload.status.code!=9 | ||
| labelExtractors: |
There was a problem hiding this comment.
Step timeouts produce status.code=9 (FAILURE), not 4 (TIMEOUT): When an individual step hits its step timeout: (for example smoke-gate hitting 1200s), Cloud Build kills the step container immediately (so RELEASE_RESULT is never printed and alert_policy.yaml does not fire) and marks the overall build as FAILURE (protoPayload.status.code=9). Because this filter has protoPayload.status.code!=9, a step timeout won't trigger either alert policy. We should either wrap the smoke-gate body in an inner timeout 1140s ... || fail_release smoke_gate "smoke gate timed out" so RELEASE_RESULT is always logged, or drop protoPayload.status.code!=9 here.
There was a problem hiding this comment.
Valid, I removed code!=9, so every failed run now sends an email, including step timeouts.
What this does
Adds a weekly release pipeline,
.ci/release.cloudbuild.yaml. It tests the image that the Build trigger already pushed for a commit, then deploys it to GKE and Cloud Run. It builds nothing. If no image exists for the commit, the run stops.Steps
vYYYY.MM.DD.<short-sha>. Stops if another release run is active.verify_release_smoke.pychecks coverage, errors, and scores, and retries one time on model failures.client_sql_smoke.pysends fixed SQL with thenoopgenerator and EvalConfig resources, in batch and streaming mode. Golden SQL must score 100, and wrong SQL must score 0.ERROR, orCRITICALline in the server log.:latestafter GKE passes.Notifications
Each run prints one
RELEASE_RESULTline with the status, stage, tag, PR list, and links.alert_policy.yamldefines a log-based alert that sends this line by email. A timeout or a cancel prints no line.The file has the placeholders
TRIGGER_IDandCHANNEL_NAME, so the repo has no project resource IDs. The live policy has the real values.Design notes
eval_clientexits 0 for any completed run, so the verifiers grade the CSV reports.Testing
In a local container with this branch's
.ci/mounted:[PASS] clean run: 6 prompts, 5 scorers[PASS]in batch and in streaming modeSetup outside this PR
weekly-evalbench-release(us-central1)mainafter the merge