Skip to content

feat(agent): send the CPU, memory and disk use of each site, and conform to contract v0.5.0 - #50

Merged
nabil1440 merged 4 commits into
agent/44-docker-statusfrom
agent/45-sites
Sep 28, 2026
Merged

nabil1440 merged 4 commits into
agent/44-docker-statusfrom
agent/45-sites

Conversation

@nabil1440

@nabil1440 nabil1440 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Closes #45.

Part of the monitoring agent stack. This PR is on top of #49. Contract v0.5.0, "sites". After this PR, the agent conforms to contract v0.5.0, and the README pins it.

What changes

  • New optional sample field sites: one item for each Docker Compose project in the home folder of the server user.
    • The agent lists the running containers (GET /containers/json), and groups them by the label com.docker.compose.project.working_dir.
    • It keeps a group when the parent of the folder is the home folder. directory is the last part of the folder, for example example.com or .fly.
  • CPU: the usage_usec deltas of the cgroup v2 cpu.stat of the containers, as a share of all the CPUs.
    • Only a container that both ticks saw, with the same id and the same cgroup, counts. A container that started or restarted in the minute is left out. docker restart keeps the id, but makes a new cgroup.
    • The time of the minute is the time between the two reads of the cgroups.
  • Memory: memory.current minus inactive_file, at the tick. Both the systemd and the cgroupfs cgroup drivers are supported.
  • Disk: at most one time each hour, a background walk measures each project folder.
    • It does not follow a symbolic link, and does not go into another file system. It counts a file with more than one hard link one time.
    • It skips what it cannot read, and writes a warning one time for each directory. The later walks log it at debug level. On a FlyWP server, the databases in ~/.fly/database belong to the container user, so .fly always has such files.
    • It runs on a thread with the idle I/O priority.
    • The next sample carries the result one time. The other samples have null.
    • Each folder has a limit of 5 minutes. A walk that hangs does not stop the next walks.
  • null and []. sites is null when the agent cannot read Docker, on cgroup v1, or without a home folder. It is [] when Docker runs and no project matches. The queue on disk keeps the two apart.
  • The agent keeps at most 1000 items, and drops an item whose directory is empty or longer than 255 characters.
  • The pin in the README, the package documentation and wire moves to v0.5.0.
  • A short first minute (see feat(agent): read the server each 10 seconds and send the peaks of the minute #46) also reads the containers, so the next sample has the CPU of each site.

Cost

  • With many sites, the queue on disk grows: about 100 bytes for each site in each sample. For 100 sites and a full queue of 24 hours (1440 samples), that is about 14 MB.

Tests

  • A fake Docker Engine and a fake cgroup tree: the grouping and the home rule, CPU and memory, a container with a new id, a container that restarted with the same id, the cgroupfs driver, the cases of null and [], and / as the home.
  • The disk: sent one time, a walk that fails, a walk that hangs, a sample that fails before the result goes, symbolic links, hard links, and a folder that cannot be read.
  • make check passes.

@nabil1440
nabil1440 added this pull request to stack #38 September 24, 2026 06:06
@nabil1440 nabil1440 changed the title agent/45 sites feat(agent): send the CPU, memory and disk use of each site, and conform to contract v0.5.0 Sep 24, 2026
@nabil1440
nabil1440 marked this pull request as ready for review September 24, 2026 06:07
…orm to contract v0.5.0

Contract v0.5.0, "sites": one item for each Docker Compose project in
the home folder of the server user. The CPU comes from the cgroup v2
usage_usec of the containers that both ticks saw, as a share of all the
CPUs. The memory is memory.current minus inactive_file. A background
walk measures the disk use of each folder at most each hour, at the idle
I/O priority, and the next sample carries it one time.

sites is null when the agent cannot read Docker or on cgroup v1, and []
when Docker runs and no project matches.

The README pins contract v0.5.0.

Closes #45
…d containers

- The sites come after each step of Sample that can fail, so a dropped
  sample does not take the hourly disk result with it.
- Each folder walk runs in a goroutine of its own: a walk that hangs in a
  system call no longer stops the next walks after the 5 minute limit.
- The CPU time of the containers uses the time of their own reads, not
  the time of the tick reading.
- A container that restarted with the same id has a new cgroup: its CPU
  time is left out for that minute.
- A home folder of / makes no site.
… be read

On a FlyWP server the databases in ~/.fly belong to the container user,
so the walk skips them each hour. The agent now warns one time for each
directory in a process, and logs the later walks at debug level.
A first minute without a sample now also reads the containers, so the
next sample has the CPU of each site, not null.
@nabil1440
nabil1440 merged commit 9e34c2b into develop Sep 28, 2026
3 checks passed
@nabil1440
nabil1440 deleted the agent/45-sites branch September 28, 2026 03:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feature: Send the CPU, memory and disk use of each site

1 participant