Skip to content

fix(soc-optimization-unified): make the health dashboard distinguish live from historical - #1177

Merged
scottbrumley merged 1 commit into
mainfrom
fix/health-dashboard-recency
Sep 24, 2026
Merged

scottbrumley merged 1 commit into
mainfrom
fix/health-dashboard-recency

Conversation

@scottbrumley

Copy link
Copy Markdown
Contributor

The headline tiles were 30-day aggregates, so a DC opening the dashboard today reads a 60% command error rate and starts firefighting something that resolved five days ago.

Same query, three windows, measured on a reference tenant:

window error rate
30 days 60% 3,830 of 6,335
7 days 5% 127 of 2,626
24 hours 0 failures

Changes

  • new Failures in the Last 24h tile, first in the row — the live signal
  • Command Error Rate, Failed Commands and Minutes Credited to Failed Commands move to a 7-day window, labelled (7d)
  • the trend, breakdowns and drilldown stay at 30 days, which is what they are for

A high 7-day rate beside a zero 24h count now reads as already resolved rather than on fire.

Why

That distinction was not available before, and its absence produced two false escalations in a single session:

  • the 66% command error rate turned out to be one bad day, with zero failures since
  • socfw-post-to-dataset failing 1,274 times turned out to run 1–19 Sep and stop; the task responsible for 935 of them exists in neither the repo nor the tenant, so it went away with a superseded playbook. In the last three days: 15,589 command tasks, 2 errors.

Both were raised as P0 off an undated aggregate. The dashboard was reporting truthfully and being read wrongly, which is a dashboard problem.

Verification

Every query run against a reference tenant at each of the three windows before it went in the file. check_contribution including the upload step, green.

…live from historical

The headline tiles were 30-day aggregates, so a DC opening the dashboard today
read a 60% command error rate and would start firefighting something that
resolved five days ago.

Measured on a reference tenant, same query, three windows:

  30 days   60%  (3,830 of 6,335)
   7 days    5%  (127 of 2,626)
  24 hours   0 failures

Changes:
- new 'Failures in the Last 24h' tile, first in the row - the live signal
- Command Error Rate, Failed Commands and Minutes Credited to Failed Commands
  move to a 7-day window and are labelled (7d)
- the trend, breakdowns and drilldown stay at 30 days, which is what they are
  for

A high 7-day rate beside a zero 24h count now reads as 'already resolved'
rather than 'on fire'. That distinction was not available before, and its
absence produced two false escalations in one session - the 66% error rate
and the socfw-post-to-dataset write failures were both single-period events
already over by the time they were found.
@scottbrumley scottbrumley added the version:patch Bug fix or hotfix → x.x.N label Sep 24, 2026
@scottbrumley
scottbrumley merged commit 4229624 into main Sep 24, 2026
14 of 24 checks passed
@scottbrumley
scottbrumley deleted the fix/health-dashboard-recency branch September 24, 2026 22:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

version:patch Bug fix or hotfix → x.x.N

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant