Skip to content

feat(agent): keep samples and events on disk until the control plane accepts them - #32

Merged
nabil1440 merged 3 commits into
agent/26-runfrom
agent/27-outbox
Sep 28, 2026
Merged

nabil1440 merged 3 commits into
agent/26-runfrom
agent/27-outbox

Conversation

@nabil1440

@nabil1440 nabil1440 commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Layer 2 of 7 of the monitoring agent (#31 → #37).

Summary

  • This PR adds a queue on the disk for the samples and the events.
  • The agent removes data only when the control plane accepts it. If the control plane is not available, the agent keeps the data for up to 24 hours and sends it later.
  • Each event keeps its id, so the control plane ignores an event that it already has.
  • The agent examines each value before it sends it: one bad value can make the control plane refuse the whole request.
  • The agent obeys the control plane when it tells the agent to wait. At start, it sends agent.started at once.
  • A recovery is logged. After a failure that keeps the data (a 5xx or a network error, a 401 or a 429), the first request that succeeds logs one Info line: "the control plane accepts the requests again", with failing_for. The events and the samples log apart.

Change

  • internal/agent/outbox.go: samples.json (at most 1440) and events.json.
  • internal/agent/send.go: events first, then samples with the status; at most 100 events and 240 samples in each request; the reply rules of the contract (200, 400, 401, 422, 429, 5xx); a wait of 1, 2, 4, 8, then 10 minutes after failures.
  • internal/agent/clean.go: the value ranges of the contract.
  • internal/agent/wire: the JSON bodies of the contract.

Adversarial review

The first version was refuted. These items are fixed in this PR:

  • Integers are capped at the signed 64-bit limit. A larger value made the control plane answer 500, and the agent resent it for 24 hours.
  • Events and metrics wait on their own after a failure (maintainer decision). A broken events route no longer stops the metrics.
  • A wait counts from the start of the report, so a 5 minute wait ends 5 minutes later, also with a slow control plane.
  • A retry comes at the next tick that its wait allows, not after the next full report interval.
  • When the time of a report runs out, the rest goes with the next report, with no backoff.
  • The log shows the start of an error reply (for example the validation errors of a 400), and each item that a full queue drops.
  • A queue file that cannot be read is kept as <name>.corrupt. A command id that is not a ULID is removed.

Tests

  • Under testing/synctest with a fake control plane: each reply code, the backoff steps, resends with the same id, batches, a slow control plane, metrics while events fail, and the waits.
  • End-to-end: a control plane on loopback gets agent.started with the token.

Closes #26
Closes #27

…accepts them

- Keep two queues in the state directory: samples.json (at most 1440, the
  oldest goes first) and events.json. Each change is a crash-safe write.
- Make each event id (a ULID) when the event goes into the queue, so a
  resend has the same id. Keep only the last agent.started.
- Keep each value in the range of the contract before it goes into a queue:
  one bad value makes a 400, and a 400 drops the whole request.
- Send the events first, then the samples with the status, at most 100
  events and 240 samples in one request, oldest first.
- Follow the reply rules: drop on 200, 400 and 422 (samples); keep on 401
  (retry each 5 minutes), 429 (Retry-After), 5xx and network errors (wait
  1, 2, 4, 8, then 10 minutes). Apply report_interval from each reply.
- Send agent.started at once at start, not at the next tick.

Refs #27
…ange

Fixes from the adversarial review of this layer.

- Cap each integer at the signed 64-bit limit: the control plane (PHP)
  fails a larger value with a 500, and the agent would resend the same
  sample for 24 hours.
- Events and metrics wait on their own after a failure. A broken events
  route no longer stops the metrics. (The poll for commands, in the next
  layer, still waits until the events are sent.)
- A wait counts from the start of the report, so a 5 minute wait ends at
  the tick 5 minutes later, also when the control plane is slow.
- Samples that could not go are tried again at the next tick that their
  wait allows, not only after the next full report interval.
- When the time of a report runs out, the rest goes with the next report.
  That is not a failure of the control plane and makes no backoff.
- Log the start of an error reply, for example the validation errors of a
  400. Log each sample or event that a full queue drops.
- Keep a queue file that cannot be read as <name>.corrupt.
- Remove an event's command_id that is not a ULID, so a bad id cannot make
  a 400 that drops the other events.

Refs #27
@nabil1440
nabil1440 added this pull request to stack #38 September 22, 2026 10:40
@nabil1440 nabil1440 self-assigned this Sep 22, 2026
@nabil1440 nabil1440 changed the title agent/27 outbox feat(agent): keep samples and events on disk until the control plane accepts them Sep 22, 2026
After the warnings of a failure (a 5xx or a network error, a 401 or a
429), the first request that succeeds logs one Info line with the time
since the first failure. The events and the samples log apart.
@nabil1440
nabil1440 merged commit 6133174 into develop Sep 28, 2026
1 check passed
@nabil1440
nabil1440 deleted the agent/27-outbox branch September 28, 2026 03:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant