Skip to content

[coverage] Conformance findings: CLOUDFETCH-018 #520

Description

@peco-engineer-bot

Summary

Surfaced by the multi-language coverage fan-out while conformance-testing these SPEC-IDs against databricks/databricks-sql-nodejs. Each finding is committed as an expected-failure (xfail) test in the coverage PR — the test asserts the CORRECT (post-fix) behavior and stays red until THIS driver (databricks/databricks-sql-nodejs) is fixed, then flips green as a tripwire.

Findings

  • CLOUDFETCH-018 [thrift]: Thrift CloudFetch downloader enforces no absolute end-to-end per-chunk deadline: with every cloud GET stalled 180s the drain does not return within the 150s budget and no timeout-classified error surfaces (audit finding H05).
    • failing test: CloudFetch — Stalled Download Deadline > stalled chunk download — drain raises a timeout-classified error within an absolute per-chunk deadline [thrift] (see the coverage PR diff under tests/)
  • CLOUDFETCH-018 [sea]: SEA kernel bounds a stalled CloudFetch chunk only by a static per-request timeout inside a 5-attempt retry loop, so the two multiply instead of bounding the chunk: with every GET stalled 180s the drain does not return within the 150s budget and no timeout-classified error surfaces (audit finding H05).
    • failing test: CloudFetch — Stalled Download Deadline > stalled chunk download — drain raises a timeout-classified error within an absolute per-chunk deadline [sea] (see the coverage PR diff under tests/)
  • CLOUDFETCH-018: A stalled CloudFetch chunk download is not bounded by any absolute end-to-end per-chunk deadline on either backend: with every cloud GET held 180s the drain does not return within 150s and no error surfaces. The SEA kernel bounds each attempt only by a static per-request timeout under a 5-attempt loop ("Chunk N download failed (attempt 1/5): NetworkError … after 1 attempts … retrying"), so the two multiply instead of bounding the chunk; Thrift shows the same non-termination. Fix: enforce a wall-clock deadline over the whole chunk (connect + response headers + body + retry backoff + link refresh) and surface it as a timeout-classified terminal error (audit finding H05).

Reproduce & Expected

CLOUDFETCH-018 — A CloudFetch chunk download that STALLS -- the cloud-storage GET is accepted but response headers/body never arrive -- must be abandoned under an ABSOLUTE end-to-end wall-clock budget for that chunk,…

Reproduce:

  • Stall EVERY CloudFetch download: the proxy accepts each GET and holds it for
    180s, so no attempt ever completes and the chunk can only finish by the driver
    giving up.
  • A result large enough to be delivered via CloudFetch external links; drain it
    and expect a terminal timeout rather than a multi-minute block.

Expected (per the shared spec):

  • full assertion contract:
result:
- label: stalled_drain
  exception_thrown: true
- label: stalled_drain
  elapsed_seconds_range:
    max: 150
- label: stalled_drain
  error:
    contains:
    - timeout
    - timed out
    - deadline
protocol:
  thrift:
  - label: stalled_drain
    cloud_downloads_min: 1
  sea:
  - label: stalled_drain
    cloud_downloads_min: 1

Context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions