Skip to content

Release v0.2.0: the monitoring agent - #51

Merged
nabil1440 merged 27 commits into
mainfrom
develop
Sep 28, 2026
Merged

nabil1440 merged 27 commits into
mainfrom
develop

Conversation

@nabil1440

@nabil1440 nabil1440 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

This PR merges develop into main for the v0.2.0 release of the FlyWP server CLI (fly).

Highlights

  • A monitoring agent. fly can now run as a small background service that reports the health of the server to FlyWP. You see CPU, memory, disk, network and per-site use in the FlyWP dashboard, without FlyWP logging in over SSH to read them.
  • Safer installs and updates. Each download is checked against the published checksums before it replaces the binary.
  • Many fixes for exit codes, flags and site detection.

The monitoring agent

FlyWP installs the agent on a server as a systemd service, fly agent run. It runs as the server user, never as root. It follows the FlyWP monitoring agent contract v0.5.0.

What you get PR
The service itself. One agent for each server, and servers report at different seconds of the minute, so they do not all report at once. #31
No lost data. If FlyWP cannot be reached, the agent keeps its samples and events on disk and sends them when the connection returns. #32
Server metrics each minute: CPU, load, memory, swap, disk space, network traffic and pending OS updates. #33
Actions from the dashboard. FlyWP can restart or update the agent. Each command runs one time only, even if it arrives twice. #34
fly update keeps the agent running. After an update, the agent restarts on the new binary. #35
Automatic updates, signed. The agent installs a new release by itself once a day, only when the release is signed by a FlyWP key and at least 24 hours old. Set FLY_AGENT_AUTO_UPDATE=off to turn it off on a server. #40
Short spikes are visible. The agent reads the server every 10 seconds and reports the peak of each minute, not only the average. #46
Pressure (PSI). How long processes waited for CPU, memory or disk: a direct sign that the server is overloaded. #47
Disk activity: bytes and operations read and written, with the peak each second. #48
Docker status and CPU count on the server status page. #49
Use per site: CPU, memory and disk space of each site's containers. #50

Installs and updates

Change PR
fly update and install.sh check each download against checksums.txt, and do not install a file that does not match. #36
install.sh is more reliable: clear errors, retries, and it never lets root write in the agent's folder. #37
fly update compares versions correctly, and replaces the binary in one step, so an interrupted update cannot leave a broken fly. #22
Static binaries built with Go 1.27, so fly runs on any supported Linux. #17

Fixes

Fix PR
Correct exit codes, and errors go to stderr, so scripts can tell success from failure. #18
The site search stops at /, and --domain is validated. #19
Flags pass through to wp, exec and logs. #20
A clear warning when Docker, Docker Compose or the Docker daemon is not available. #21
OpenLiteSpeed sites are detected. #1

For maintainers

Change PR
A Makefile for build, test, lint and release. #23
CI runs make check on every change, and releases are published with gh. #25
A short README, and reference docs in docs/: the monitoring agent, development, releasing and the design decisions. #52

Checks before release

  • Tested on a real server. This code is the same as the dev build v0.2.0-dev.10544b5, which passed an end-to-end test against FlyWP.
  • CI passes on develop.
  • Reproducible build. Building the tag again gives the same binaries, byte for byte. A maintainer rebuilds the release on their own computer, compares it with the files on GitHub, and only then signs it. Agents install only signed releases by themselves.

Merge with a merge commit, not a squash: main has one commit that develop lacks.

ma-04 and others added 27 commits June 16, 2025 15:40
* fix(ols): cli not detecting ols site

* fix(ols): set ols user to avoid using root user
Install the gh-stack skill from github/gh-stack at v0.1.1 with
gh skill install github/gh-stack gh-stack --agent claude-code --scope project --pin v0.1.1
- go.mod: go 1.27.0, toolchain go1.27.1; update dependencies; go mod tidy
- build.sh: CGO_ENABLED=0, -trimpath, -s -w (static, stripped binaries)
- build.yml: go-version '1.27' with check-latest

Release asset names and tarball layout are unchanged, so the v0.1.1
updater can still install new releases.

Closes #6
* fix(cli): return correct exit codes and report errors on stderr

- Use RunE on every command and let Execute report the error once.
- Pass the exit status of docker compose and wp-cli through unchanged.
- Print other errors on stderr with a non-zero exit status.
- Remove the color.Red printf misuse that printed %!(EXTRA ...).
- Show success messages only after success.
- sites start/stop/restart continue after a failed site and report
  every failed site.
- Add a fake docker test harness and end-to-end tests.

Closes #2

* fix(cli): name the phase of each failed site in sites commands

sites restart printed 'site: exit status 1' twice when a site failed
to stop and then to start. Each failure now starts with its phase:
'stopping site: …' or 'starting site: …'.

Part of #2
- FindComposeFile stops at the home directory or at the filesystem
  root, so a command outside $HOME no longer loops forever.
- FindComposeFile returns errors from UserHomeDir and Getwd, and
  ErrComposeNotFound when a site has no compose file.
- --domain must be a valid hostname (labels of letters, digits and
  inner hyphens). This rejects path separators and '..', so the value
  cannot point outside the sites directory.

Closes #3
- wp and exec stop parsing fly flags at the first argument, so
  'fly wp plugin list --format=json' and 'fly exec php ls -la' work.
  --domain goes before the command.
- exec uses the site's PHP service by default: php, or openlitespeed
  on OpenLiteSpeed sites (DefaultService, shared with wp).
- exec reports an error when only a service name is given.
- logs accepts --follow/-f, --tail and several services.
- README: document flag passthrough and --domain placement.

Closes #4
#21)

- docker.Check examines the Docker CLI (PATH), the Compose plugin
  (docker compose version) and the daemon (docker version), with a
  5 s timeout for each check.
- Commands that need Docker declare it with requireDocker. The root
  command runs the check first. If a part is not available, fly shows
  one warning line and exits with status 69 (EX_UNAVAILABLE).
- fly status shows the three parts on separate lines and exits 0.

Closes #5
…#22)

- Compare versions with golang.org/x/mod/semver (v0.1.10 is newer
  than v0.1.9). A git describe build (v0.1.1-2-gabc) counts as its
  tag. A build that is not from a release tag (dev) cannot be
  compared: fly update says so and asks before it installs.
- Get the release one time, with a 60 s timeout, a User-Agent and a
  status check. A GitHub rate limit gives a clear error instead of
  'already running the latest version'.
- Extract the binary with archive/tar and compress/gzip instead of
  the tar command.
- Write the new binary to a temporary file next to the executable and
  rename it, so the rename cannot cross filesystems (EXDEV). The old
  binary stays unchanged when the update fails.
- Release asset names and archive layout are unchanged.

Closes #8
…23)

- make build, test, vet, lint, vuln, fmt, fmt-check, check, release,
  clean and help, in the style of migration-agent.
- lint and vuln run pinned tool versions through go run, so the tools
  are built with the Go version of this module.
- make check is the local gate before each merge (CI is skipped, #7).
- make release makes the same archives as build.sh before, plus
  build/checksums.txt. build.sh is now a wrapper for make release, so
  the release workflow does not change.
- README: document the targets.

Closes #16
* ci: run make check on every change and publish releases with gh

- ci.yml: run make check and make release on every pull request and
  on pushes to develop and main. GitHub Actions are free for this
  public repository.
- build.yml (release): one job on actions/checkout@v7 and
  actions/setup-go@v7. It checks that the tag is on main, runs
  make check, builds with make release VERSION=<tag> and publishes
  with gh release create: both archives plus checksums.txt, titled
  with the tag. A tag with a pre-release suffix becomes a pre-release,
  so installed CLIs do not update to it.
- Remove the archived create-release and upload-release-asset actions
  and the build.sh wrapper, which nothing calls now.
- README: document CI and the release steps.

Asset names and the archive layout are unchanged.

Closes #7

* ci: publish dev pre-releases (vX.Y.0-dev.<sha>) from any branch

- make dev-version prints <next minor>-dev.<short commit>, for example
  v0.2.0-dev.1a2b3c4. DEV_BASE overrides the version part.
- make dev-release tags HEAD with it and pushes the tag. It refuses
  uncommitted changes, commits that are not pushed, and commits whose
  release workflow would publish the tag as a full release.
- The Release workflow allows pre-release tags from any branch. It
  publishes them as pre-releases, so /releases/latest (and so
  fly update and install.sh) never returns them. Release tags without
  a suffix must still be on main.
- README: document dev pre-releases and how to install one on a
  test server.

Part of #7
chore: add gh-stack agent skill
* feat(agent): add fly agent run, the base of the monitoring agent

- Read FLY_AGENT_URL, FLY_AGENT_TOKEN, FLY_AGENT_SERVER_ID and
  STATE_DIRECTORY. An error names each key that is not set or not valid.
- Accept only an https URL. Plain http is accepted only for a loopback
  host, for tests and local development.
- Lock a file in the state directory, so that only one agent operates.
- Work at server_id % 60 seconds past each minute, and report after each
  report interval (1 to 10 samples, saved in state.json).
- Stop on SIGTERM. The agent does not need Docker and does not run as root.
- Add a client for the control plane and internal/statefile for crash-safe
  JSON files. The next layers use them.

Refs #26

* fix(agent): refuse redirects, cap Retry-After, and never tick two times

Fixes from the adversarial review of this layer.

- The control plane client does not follow redirects. A followed redirect
  replays a POST as a GET without the body, so a 200 to it would drop
  samples that were never stored, and a redirect to http would send the
  token in plain text. A 3xx is now a reply that keeps the data.
- A Retry-After value is at most one hour and cannot overflow.
- A wall clock that steps back cannot make the same tick run two times.
- FLY_AGENT_URL must not hold a user, a password, a query or a fragment,
  and an error never shows a password. localhost is accepted in any case.
- FLY_AGENT_TOKEN may hold only printable ASCII characters.
- At start, remove the temporary files that a crash during a write left.

Refs #26

* test(docker): give the probe timeout test room under load

Under go test ./... -race, the fake "docker compose version" could take
more than the 1 s test timeout, so the test reported the Compose plugin
instead of the daemon. Use 3 s.

Refs #26
…accepts them (#32)

* feat(agent): keep samples and events on disk until the control plane accepts them

- Keep two queues in the state directory: samples.json (at most 1440, the
  oldest goes first) and events.json. Each change is a crash-safe write.
- Make each event id (a ULID) when the event goes into the queue, so a
  resend has the same id. Keep only the last agent.started.
- Keep each value in the range of the contract before it goes into a queue:
  one bad value makes a 400, and a 400 drops the whole request.
- Send the events first, then the samples with the status, at most 100
  events and 240 samples in one request, oldest first.
- Follow the reply rules: drop on 200, 400 and 422 (samples); keep on 401
  (retry each 5 minutes), 429 (Retry-After), 5xx and network errors (wait
  1, 2, 4, 8, then 10 minutes). Apply report_interval from each reply.
- Send agent.started at once at start, not at the next tick.

Refs #27

* fix(agent): send metrics while events fail, and keep every value in range

Fixes from the adversarial review of this layer.

- Cap each integer at the signed 64-bit limit: the control plane (PHP)
  fails a larger value with a 500, and the agent would resend the same
  sample for 24 hours.
- Events and metrics wait on their own after a failure. A broken events
  route no longer stops the metrics. (The poll for commands, in the next
  layer, still waits until the events are sent.)
- A wait counts from the start of the report, so a 5 minute wait ends at
  the tick 5 minutes later, also when the control plane is slow.
- Samples that could not go are tried again at the next tick that their
  wait allows, not only after the next full report interval.
- When the time of a report runs out, the rest goes with the next report.
  That is not a failure of the control plane and makes no backoff.
- Log the start of an error reply, for example the validation errors of a
  400. Log each sample or event that a full queue drops.
- Keep a queue file that cannot be read as <name>.corrupt.
- Remove an event's command_id that is not a ULID, so a bad id cannot make
  a 400 that drops the other events.

Refs #27

* feat(agent): log when the control plane accepts the requests again

After the warnings of a failure (a 5xx or a network error, a 401 or a
429), the first request that succeeds logs one Info line with the time
since the first failure. The events and the samples log apart.
* feat(agent): measure the server each minute

- Measure CPU, load, memory, swap, disk and network from /proc and
  statfs("/"). Used memory is MemTotal - MemAvailable; used disk is
  (blocks - free blocks), as the contract defines them.
- Count only the network interfaces that have a hardware device, so that
  container traffic through veth, the bridges and docker0 is not counted two
  or three times. Without such an interface, count the default route.
- Save the last counters with the boot id. After an agent restart the next
  sample continues from them. After a reboot, a counter that went back or a
  reading older than 90 seconds, send 0 with net_counters_reset.
- Send the status in each report: reboot required, the update counts from
  apt-check (each hour; 0 and a warning when it fails), the OS name, the
  kernel, the uptime and the arch.
- A measurement that fails skips that minute. The agent continues.

Refs #28

* fix(agent): read the update counts of apt-check correctly on Ubuntu 24.04

Fixes from the adversarial review of this layer.

- apt-check can write warnings before its result, for example for a source
  that is configured two times. Read only the last line. Before, both counts
  became 0, and a server with security updates looked up to date.
- A failed count keeps the last counts. The counts are 0 only when apt-check
  never gave a result.
- Count the updates with the sample of the minute, before the sends of a
  report, so a slow apt-check cannot use the time of the sends.
- When no interface is counted, mark the traffic as not known
  (net_counters_reset), not as 0, and log it one time.
- The default route must have mask 0, so a VPN route 0.0.0.0/1 is not taken.
- Leave out an interface that is a port of an other interface (a bond or a
  bridge port, or the Azure VF under netvsc): its traffic is also in the
  interface above it.

Refs #28
* feat(agent): run agent.update and agent.restart one time only

- Poll GET /agent/v1/commands after each report, and run the new commands
  one at a time, oldest first. Send the results at once.
- Keep a list of the commands that ran (ran.json, 48 hours) and skip them:
  the poll sends each open command again until its result arrives.
- agent.restart: record the command, then exit; systemd starts the agent
  again, and the new process sends command.completed.
- agent.update: skip the download when the agent already runs the target
  (or a newer release). Else download the archive next to the binary,
  compare its sha256, and put the new binary in place with a rename. A
  failure sends command.failed and keeps the old binary. After the exit,
  the new process compares its version with the target and sends the result.
- A dev tag (v0.2.0-dev.<sha>) matches only the same tag: dev tags have no
  order.
- Do not run an unknown verb; send command.unknown.
- Move the update code of fly update to internal/release, and add Download
  and Install for the agent.
- README: the contract pin line and the monitoring agent section.

Refs #29

* fix(agent): install an older release on request, and write the binary safely

Fixes from the adversarial review of this layer.

- agent.update skips the download only when the agent runs exactly the
  target version. Before, an older target reported command.completed and
  installed nothing, so a rollback looked done. The control plane decides
  the version; after the update, a newer release still completes it.
- An update download follows a redirect only to https (or loopback).
- At start, remove downloads that a crash left next to the binary, when
  they are older than one hour.
- Make the new binary executable through the open file, not through its
  path: the directory can belong to an other user. Sync the binary and the
  directory, so a power loss cannot leave an empty binary.

Refs #29

* fix(agent): never downgrade on agent.update

Maintainer decision: when the agent already runs the target version or a
newer one, it does not update.

- The same version sends command.completed, as before.
- An older target sends command.failed with the reason ("the agent runs
  v0.2.1, newer than v0.2.0; it does not downgrade"), so the control plane
  sees that its version was not installed. A bad release is fixed with a
  newer release; an older version needs the install job.
- Versions compare with semver. Two dev tags of one release
  (v0.2.0-dev.<sha>) have no order, and a version that is not semver has no
  order: then only the exact version skips the update.
- The new process after an update uses the same comparison for its result.

Refs #29
* feat(update): keep the monitoring agent correct after fly update

- After fly update replaces the binary, restart fly-agent when the server
  has the agent. The agent then runs the new binary at once.
- When fly is already up to date but the agent still runs a replaced binary
  (the kernel shows it as "(deleted)"), restart the agent.
- Keep the owner of the old binary: sudo fly update no longer gives the
  agent's binary in ~fly/.fly/bin to root.
- Without --yes and without a terminal, stop with an error (exit 1). A
  script no longer reads "Update cancelled." as success.
- Add internal/service for the systemd unit of the agent.

Refs #30

* fix(update): change the owner through the open file, and show systemctl errors

Fixes from the adversarial review of this layer.

- Give the new binary its owner through the open file (fchown), not its
  path. The folder of the agent's binary belongs to the server user, who
  could put a link to a root file (for example /etc/shadow) in place of
  the temporary file, and root would then give that file to the user.
- A systemctl failure now shows its message. Before, fly update exited 1
  with no message, because the wrapped exit status looked like the error
  of a child process that had already reported it.
- Restart the agent with try-restart: an agent that an administrator
  stopped stays stopped.
- An agent process that ends during the check is not stale.
- Wait up to 2 minutes for systemctl, longer than the stop timeout of
  systemd.
- CI runs the tests that need root (the owner and the link attack).

Refs #30
* feat(update): examine release downloads with checksums.txt

- fly update downloads checksums.txt of the release and compares the
  sha256 of the archive before it extracts the binary. A release without
  checksums.txt, or without a line for the archive, is not installed.
- fly update now uses the same download and check code as agent.update,
  with a limit of 10 minutes.
- install.sh downloads checksums.txt and runs sha256sum -c. It stops when
  the file is missing or the checksum does not agree.

Refs #9

* fix(update): refuse two sums for one file, and stop the download on Ctrl-C

Fixes from the adversarial review of this layer.

- checksums.txt with two different sums for the same file is not valid.
  Before, fly update used the first line and install.sh refused.
- Ctrl-C or SIGTERM during fly update stops the download and removes the
  partial file.
- README: releases before v0.2.0 have no checksums.txt, so push the tag
  right after the merge into main.
- Tests for CRLF line ends, upper case hex, repeated and similar lines, an
  empty checksum file, and no files left after a successful update.

Refs #9
…he agent folder (#37)

* fix(install): make install.sh reliable, and install through the agent link

- Use set -euo pipefail, curl -fsSL and a trap that removes the temporary
  files. Report the HTTP status of the GitHub API, with a rate-limit message
  for 403 and 429. Parse tag_name with sed; remove the grep -P step.
- Extract only fly-<os>-<arch>, without the owner from the archive, and
  install it as root:root. Before, the binary in /usr/local/bin belonged to
  uid 1001 (the CI user), so a local user with that uid could replace a
  binary that root runs.
- On a server with the monitoring agent, /usr/local/bin/fly is a link to
  ~fly/.fly/bin/fly: install through the link, keep the owner of the
  target, and restart fly-agent. The CLI and the agent keep one binary.
- Install with a rename, so a running fly never sees a partial binary.
- make release: the archives hold the binary as root:root (GNU tar and
  bsdtar).

Refs #10

* fix(install): never let root write in the folder of the agent user

Fixes from the adversarial review of this layer.

- On a server with the monitoring agent, the agent user writes the new
  binary (runuser), and root only gives it the file on stdin. Before, root
  followed /usr/local/bin/fly into ~fly/.fly/bin: the fly user could point
  the binary at /etc/passwd and make root overwrite it, or swap in a link
  during the copy and make root give it /etc/shadow.
- Take the user and the binary from fly-agent.service (User=, ExecStart=)
  and require <home>/.fly/bin/fly. A server that the old installer split
  (a separate /usr/local/bin/fly) gets one binary again, with the link.
- Without the agent, install a root file with a temporary name and a
  rename. A link at /usr/local/bin/fly is replaced, never followed, and a
  leftover fly.new cannot break the install.
- Restart the agent with try-restart, so a stopped agent stays stopped.
- Read checksums.txt with the same rules as fly update ("*" mark, CRLF).
- Take the first tag_name of the one-line API reply. Add timeouts and
  retries to each download. The install command uses curl -fsSL.

Refs #10
* feat(release): sign checksums.txt with a key kept outside GitHub

Add the signature file checksums.txt.sig: the key id, the tag, the time of
the signature and an ed25519 signature over these lines and the bytes of
checksums.txt. SignedChecksum checks it and takes the sum from the same
bytes. The tool tools/releasesign makes the key and signs a release; make
release-key and make sign-release run it. keys.go holds the trusted
public keys.

Refs #39

* feat(agent): install a new signed release by itself, one time each day

Each day, at a time from the server id, the agent gets the latest release.
It installs the release only when it is newer, its checksums.txt has a valid
signature, and the signature is more than 24 hours old. It then exits, and
systemd starts the new binary. The check runs on each tick, not only on a
report tick, and it saves its time before it starts.

FLY_AGENT_AUTO_UPDATE=off stops the check on one server. A build without a
release version, or without a trusted key, does not check. agent.update
works as before.

The state file now keeps the whole state: a new report interval no longer
removes the time of the last check.

Closes #39

* fix(release): sign only a release that builds again from the local tag

make sign-release signed whatever checksums.txt GitHub served. Someone who
controls GitHub could swap an archive and its sum before the signature.

tools/sign-release.sh now shows the commit of the local tag and asks to
type the tag. It checks the archives against checksums.txt, builds the
release again from the local tag with the Go version of the CI binaries,
and compares the binaries byte for byte. The binary holds its commit, so
this also proves that CI built the local tag. checksums.txt must name
exactly the built archives. The script signs and verifies with the tool
and the keys of the tag.

The build date is now the date of the commit, so the same commit and Go
version give the same binary on any computer.

Refs #39

* fix(agent): no auto-update for local builds, and limit the release reply

- A make build after a tag (v0.2.0-3-gabcdef1, -dirty) sorts before the
  tag, so the release would replace newer code. release.IsLocalBuild
  turns auto-update off for these builds.
- Read the release JSON through a 1 MiB limit: the agent reads it each
  day without a person.
- Test that the tag and the key lines are part of the signed bytes.

Refs #39

* fix(agent): follow the clock after a large step back

After the wall clock stepped back by more than a minute, the loop waited
for the old tick time: the agent sent nothing for the size of the step.
For example, a VM that booted with its clock one hour ahead went silent
for one hour after NTP corrected it.

The loop now skips a repeated minute only after a step back of less than
2 minutes. After a larger step it follows the new clock. A minute that
runs two times is harmless: the control plane keeps one sample for each
minute.

* feat(release): name the release key with a comment

make release-key takes COMMENT (default: "server-cli release key for
flywp"). The comment goes in the private key file as a PEM header and in
the output, and keys.go names each key with it. A comment must be one
line, so that it cannot add an other PEM header.

Refs #39

* fix(release): accept ~/ in the key path of release-key and sign-release

In "make release-key KEY=~/key" the shell does not expand the "~": it is
not at the start of a word. The tool and the sign script now replace a
leading "~/" with the home directory.

Refs #39

* feat(release): trust the server-cli release key for flywp

Add the public key 21ec14c790e96d62 ("server-cli release key for flywp").
The private key is kept outside GitHub. A test now fails when the list of
trusted keys is empty, so that no release ships without a key.

Refs #39

* feat(agent): conform to contract v0.3.0, with null for unknown update counts

Contract v0.3.0 records the update by release, and permits null for the
update counts. The agent now sends updates_total and updates_security as
null until apt-check gives a count. Before, it sent 0: a false "no
updates". A later failure keeps the last counts, as before.

The README pin is v0.3.0.

Refs #39

* docs(agent): conform to contract v0.3.1

The contract now states that a failed update count keeps the last counts,
and that null means no count since the agent started. The agent already
does this; only the pin changes.
…e minute (#46)

* feat(agent): read the server each 10 seconds and send the peaks of the minute

Contract v0.4.0, "The readings" and "The peaks": cpu_max_percent,
memory_used_max_bytes, swap_used_max_bytes and the two network peaks.
The fields of v0.3.1 do not change. apt-check now runs after the reading
of the tick, so that it does not move the reading.

Closes #41

* fix(agent): keep the tick after a slow reading, and time the windows by the real reading times

- A reading that ends at or after the next tick no longer skips it.
- Each reading, and the tick reading, carries the time at which it ran
  (with the monotonic clock), so a late timer or a step of the wall clock
  does not change the length of a window. recorded_at stays the tick.
- cpu_max_percent leaves out the windows of an older minute whose tick
  failed, as the memory peaks do.

* fix(agent): send no sample for a first minute shorter than 10 seconds

After a fresh start (no saved counters), the first tick can come some
milliseconds after the start reading: its CPU value is the load of the
start, for example 100% during an install. That minute now has no
sample, and the tick reading starts the next minute, which then has all
its values. The agent logs it at Info, not as a warning.

* test(agent): correct a comment about the reports without a sample

* test(agent): measure the first minute in the test on a real Linux server
…isk (#47)

* feat(agent): send the pressure (PSI) of the CPU, the memory and the disk

Contract v0.4.0, "Pressure (PSI)": the share of the minute in which at
least one task waited, from the "some" total of /proc/pressure, and
its peak within the minute. null without PSI, in the first sample, after
a reboot and after a counter went back.

Closes #42

* fix(agent): join the pressure windows around a reading without PSI
)

* feat(agent): send the disk activity, and conform to contract v0.4.0

Contract v0.4.0, "Disk activity": the bytes and the operations that the
hardware disks read and wrote in the minute, from /proc/diskstats, and
their peaks each second within the minute. null in the first sample,
after a reboot, after a counter went back, and without a hardware disk.

The README pins contract v0.4.0.

Closes #43

* fix(agent): send no disk activity when a counter went back within the minute

The contract makes all eight disk fields null when a counter went back.
A disk that is attached again under the same name can pass the check of
the minute and still give a wrong delta; now a window that went back
nulls the values of the minute too. A reading with no disks for a moment
joins its windows.
Contract v0.5.0, "status: the CPU count and Docker": cpu_count from the
cpuN lines of /proc/stat, and docker_status and docker_version from
GET /version on the Docker socket, with a limit of 5 seconds. The new
package dockerapi uses only the Go standard library, and sends only
GET /version and GET /containers/json.

Closes #44
…orm to contract v0.5.0 (#50)

* feat(agent): send the CPU, memory and disk use of each site, and conform to contract v0.5.0

Contract v0.5.0, "sites": one item for each Docker Compose project in
the home folder of the server user. The CPU comes from the cgroup v2
usage_usec of the containers that both ticks saw, as a share of all the
CPUs. The memory is memory.current minus inactive_file. A background
walk measures the disk use of each folder at most each hour, at the idle
I/O priority, and the next sample carries it one time.

sites is null when the agent cannot read Docker or on cgroup v1, and []
when Docker runs and no project matches.

The README pins contract v0.5.0.

Closes #45

* fix(agent): keep the disk result of the sites, and leave out restarted containers

- The sites come after each step of Sample that can fail, so a dropped
  sample does not take the hourly disk result with it.
- Each folder walk runs in a goroutine of its own: a walk that hangs in a
  system call no longer stops the next walks after the 5 minute limit.
- The CPU time of the containers uses the time of their own reads, not
  the time of the tick reading.
- A container that restarted with the same id has a new cgroup: its CPU
  time is left out for that minute.
- A home folder of / makes no site.

* fix(agent): warn one time for each site folder with files that cannot be read

On a FlyWP server the databases in ~/.fly belong to the container user,
so the walk skips them each hour. The agent now warns one time for each
directory in a process, and logs the later walks at debug level.

* fix(agent): start the CPU times of the sites at a short first minute

A first minute without a sample now also reads the containers, so the
next sample has the CPU of each site, not null.
The README is short: install, the commands, the agent in one paragraph,
and links. The details move to docs/: the monitoring agent, development,
releasing and the design decisions.
docs: rewrite the README and add the reference docs of v0.2.0
@nabil1440
nabil1440 merged commit 1e29ec0 into main Sep 28, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants