Skip to content

Make desktop E2E results a trustworthy merge signal #417

Description

@kostyafarber

Problem

Desktop E2E results are not a reliable merge signal. Over 600 CI runs from 2026-08-18 to 2026-09-25, about 6% of successful E2E jobs passed only because a test succeeded on its Playwright retry. Two tests produced most of these:

  • application-quit re-entrant quit: 18 retry passes and 1 hard failure;
  • glyph-view zoom momentum: 9 retry passes and 5 hard failures, including on pull requests that changed no code.

Both depend on fixed sleeps.

Other oracles are coupled to incidental rendering details, or can pass when the behavior is broken:

  • Duplicate layer baselines. The "layer" canvas snapshots are byte-identical composites.
  • No tolerance or stabilization on buffer snapshots. Buffer toMatchSnapshot comparisons skip Playwright's stable-capture retry and the configured tolerance.
  • Pixel-colour oracles. Several tests count hard-coded palette colours. These break under theme changes, and some would still pass with handles or outlines missing.
  • Broken perf gate. The nightly performance run has failed on every scheduled run since 2026-09-07, and its committed baseline never gates a result.

Expected outcome

A red E2E job means a product regression or an intentional visual change. Tests wait on observable application state, not elapsed time. Rendering assertions compare focused, settled captures or semantic render state instead of palette colours. Flaky results stay visible instead of being absorbed by retries.

Acceptance criteria

  • Desktop E2E specs and fixtures contain no fixed waitForTimeout delays.
  • Quit, close, and relaunch tests wait on recorded main-process decisions, and each launch starts from an environment without inherited SHIFT_E2E_* variables.
  • Every golden capture uses toHaveScreenshot on a focused locator after the canvas has rendered, and no two baselines duplicate the same image.
  • No E2E assertion counts or samples hard-coded palette colours.
  • Recovery, save, and export tests assert persisted content rather than file existence or byte inequality.
  • Relaunched Electron processes are owned and cleaned up by fixtures, including after a test timeout.
  • CI reports tests that pass only on retry, and golden visual comparisons do not retry.
  • The nightly performance workflow runs its benchmarks to completion and measures rendered frames.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions