fix(soc-optimization-unified): make the health dashboard distinguish live from historical - #1177
Merged
Merged
Conversation
…live from historical The headline tiles were 30-day aggregates, so a DC opening the dashboard today read a 60% command error rate and would start firefighting something that resolved five days ago. Measured on a reference tenant, same query, three windows: 30 days 60% (3,830 of 6,335) 7 days 5% (127 of 2,626) 24 hours 0 failures Changes: - new 'Failures in the Last 24h' tile, first in the row - the live signal - Command Error Rate, Failed Commands and Minutes Credited to Failed Commands move to a 7-day window and are labelled (7d) - the trend, breakdowns and drilldown stay at 30 days, which is what they are for A high 7-day rate beside a zero 24h count now reads as 'already resolved' rather than 'on fire'. That distinction was not available before, and its absence produced two false escalations in one session - the 66% error rate and the socfw-post-to-dataset write failures were both single-period events already over by the time they were found.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The headline tiles were 30-day aggregates, so a DC opening the dashboard today reads a 60% command error rate and starts firefighting something that resolved five days ago.
Same query, three windows, measured on a reference tenant:
Changes
(7d)A high 7-day rate beside a zero 24h count now reads as already resolved rather than on fire.
Why
That distinction was not available before, and its absence produced two false escalations in a single session:
socfw-post-to-datasetfailing 1,274 times turned out to run 1–19 Sep and stop; the task responsible for 935 of them exists in neither the repo nor the tenant, so it went away with a superseded playbook. In the last three days: 15,589 command tasks, 2 errors.Both were raised as P0 off an undated aggregate. The dashboard was reporting truthfully and being read wrongly, which is a dashboard problem.
Verification
Every query run against a reference tenant at each of the three windows before it went in the file.
check_contributionincluding the upload step, green.