Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 12 additions & 4 deletions README.MD
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,9 @@ SyMon is a self-hosted monitoring tool for Linux servers, home labs and Raspberr
**Hosts**
- CPU, overall and per core, load average, memory and swap
- Disk space, disk IO and busy time, network traffic, TCP connections
- When each disk will be full, at the rate it grew over the last week
- Pressure stall information and temperature sensors
- The top processes by CPU and memory, at any point in time
- The top processes by CPU and memory at any point in time, and the programs that used the most over any range
- Whether chosen systemd services are running

**Containers**
Expand All @@ -23,18 +24,19 @@ SyMon is a self-hosted monitoring tool for Linux servers, home labs and Raspberr
- Send any number from a script or cron job and get a chart for it

**Alerts**
- Rules for CPU, memory, swap, disks, services, custom metrics, silent hosts and HTTP endpoints
- Rules for CPU, memory, swap, disks, disks filling up, services, custom metrics, silent hosts and HTTP endpoints
- Warning and critical levels, shown on the dashboard and sent by email, Slack or PagerDuty

**Dashboard**
- Every host at a glance, and a page per host with charts from 15 minutes to 30 days, or any custom range
- Drag across a chart to zoom in, and switch any chart to a table
- Drag across a chart to zoom in, click a point to see the processes running then, and switch any chart to a table
- Follows the system light or dark theme

**Running it**
- Add a host with one command, with a single-use token
- Raw data kept for 7 days, 1 minute averages for 30 days and 1 hour averages for a year, all adjustable
- Each host has its own key, and components can talk over TLS
- A Prometheus endpoint, for Grafana or a Prometheus you already run

## Screenshots

Expand Down Expand Up @@ -96,15 +98,21 @@ Components talk over gRPC, so other tools can read from or push into them. See t
The Client exposes a JSON API under `/api/v1`. Times are unix seconds. Errors return a JSON body `{"error": "..."}` with a 4xx or 5xx status.

* `GET /api/v1/fleet`
* Every host with its status, latest usage, number of running containers and number of open alerts
* Every host with its status, latest usage, number of running containers, number of open alerts, and `diskFullDays`, the days until its first disk is full (null when none is filling up)
* `GET /api/v1/hosts/{host}`
* The host's latest snapshot as the agent sent it, with `up` and `lastSeen`
* `GET /api/v1/hosts/{host}/series?metric=cpu&from=&to=`
* A metric over time. `from` and `to` default to the last hour. `label` keeps one series (a mount point, interface, sensor, container or custom metric name), `maxPoints` sets the most points per series (default 1000) and `max=1` returns the peak of each bucket instead of the average
* Metrics: `cpu`, `cpu_core`, `load1`, `load5`, `load15`, `memory`, `memory_used`, `swap`, `swap_used`, `psi_cpu`, `psi_memory`, `psi_memory_full`, `psi_io`, `psi_io_full`, `tcp_established`, `tcp_time_wait`, `tcp_close_wait`, `tcp_listen`, `tcp_total`, `disk_used`, `disk_inodes`, `disk_read`, `disk_write`, `disk_util`, `net_rx`, `net_tx`, `temperature`, `custom`, `container_cpu`, `container_memory`, `container_rx`, `container_tx`, `container_io_read`, `container_io_write`
* `GET /api/v1/hosts/{host}/processes?at=`
* Top processes by CPU and by memory at or before `at`, or the latest
* `GET /api/v1/hosts/{host}/process-usage?from=&to=`
* The programs that used the most CPU and memory over the range, 24 hours at most, with their average, peak and how often they were among the top processes. Processes with the same name are added up
* `GET /api/v1/hosts/{host}/disk-forecasts`
* Each disk's growth per day over the last week, and `daysToFull`, or null when the disk is not filling up
* `GET /api/v1/hosts/{host}/custom-metrics`
* Names of the host's custom metrics
* `GET /api/v1/alerts?host=&open=1&from=&to=`
* Alerts, newest first. `open=1` leaves out resolved ones

The Client also serves `GET /metrics`, every host's latest values in the Prometheus text format. See [Prometheus and Grafana](docs/install.md#prometheus-and-grafana).
234 changes: 234 additions & 0 deletions client/internal/server/metrics.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,234 @@
package server

import (
"encoding/json"
"fmt"
"io"
"net/http"
"strconv"
"strings"

"github.com/dhamith93/SyMon/internal/api"
"github.com/dhamith93/SyMon/internal/logger"
"github.com/dhamith93/SyMon/internal/monitor"
)

// getMetrics serves every host's latest values in the Prometheus text
// format, so Prometheus can scrape the dashboard
func (s *server) getMetrics(w http.ResponseWriter, r *http.Request) {
response, err := s.collector.Snapshots(r.Context(), &api.Void{})
if err != nil {
writeGRPCError(w, "snapshots", err)
return
}

metrics := newMetricsWriter()
for _, host := range response.Hosts {
metrics.gauge("symon_up", "1 when the host reported within the last minute.", boolValue(host.Up), "host", host.Host)
if host.LastSeen > 0 {
metrics.gauge("symon_last_seen_timestamp_seconds", "When the host was last heard from.", float64(host.LastSeen), "host", host.Host)
}
// a host that stopped reporting only has old values, which would
// look current to Prometheus
if !host.Up || host.SnapshotJson == "" {
continue
}
var data monitor.MonitorData
if err := json.Unmarshal([]byte(host.SnapshotJson), &data); err != nil {
logger.Log("error", "cannot read the snapshot of "+host.Host+": "+err.Error())
continue
}
addHostMetrics(metrics, host.Host, &data)
}

w.Header().Set("Content-Type", "text/plain; version=0.0.4; charset=utf-8")
if err := metrics.writeTo(w); err != nil {
logger.Log("error", "cannot write metrics: "+err.Error())
}
}

const mib = 1024 * 1024

func addHostMetrics(m *metricsWriter, host string, data *monitor.MonitorData) {
m.gauge("symon_uptime_seconds", "How long the host has been running.", data.System.UpTimeSeconds, "host", host)
// LoadAvg keeps its old name, it is the CPU usage
m.gauge("symon_cpu_usage_percent", "CPU usage.", float64(data.ProcUsage.LoadAvg), "host", host)
m.gauge("symon_load1", "Load average over 1 minute.", data.ProcUsage.Load1, "host", host)
m.gauge("symon_load5", "Load average over 5 minutes.", data.ProcUsage.Load5, "host", host)
m.gauge("symon_load15", "Load average over 15 minutes.", data.ProcUsage.Load15, "host", host)

// the agent sends memory and swap in MiB
m.gauge("symon_memory_used_percent", "Memory in use.", data.Memory.PercentageUsed, "host", host)
m.gauge("symon_memory_total_bytes", "Total memory.", float64(data.Memory.Total)*mib, "host", host)
m.gauge("symon_memory_available_bytes", "Memory available to new programs.", float64(data.Memory.Available)*mib, "host", host)
m.gauge("symon_swap_used_percent", "Swap in use.", data.Swap.PercentageUsed, "host", host)
m.gauge("symon_swap_total_bytes", "Total swap.", float64(data.Swap.Total)*mib, "host", host)

for _, disk := range data.Disk {
labels := []string{"host", host, "device", disk.FileSystem, "mount", disk.MountedOn}
if pct, ok := parsePercent(disk.Usage.Usage); ok {
m.gauge("symon_disk_used_percent", "Disk space in use, as df shows it.", pct, labels...)
}
m.gauge("symon_disk_size_bytes", "Disk size.", float64(disk.Usage.Size), labels...)
m.gauge("symon_disk_used_bytes", "Disk space in use.", float64(disk.Usage.Used), labels...)
if pct, ok := parsePercent(disk.Inodes.Usage); ok {
m.gauge("symon_disk_inodes_used_percent", "Inodes in use.", pct, labels...)
}
}
for _, diskIO := range data.DiskIO {
// named like the disks above
labels := []string{"host", host, "device", "/dev/" + diskIO.Device}
m.gauge("symon_disk_read_bytes_per_second", "Disk reads.", diskIO.ReadBytesPerSec, labels...)
m.gauge("symon_disk_write_bytes_per_second", "Disk writes.", diskIO.WriteBytesPerSec, labels...)
m.gauge("symon_disk_util_percent", "Time the disk was busy.", diskIO.UtilPercent, labels...)
}
for _, network := range data.Networks {
labels := []string{"host", host, "iface", network.Interface}
m.counter("symon_network_receive_bytes_total", "Bytes received since boot.", float64(network.Usage.RxBytes), labels...)
m.counter("symon_network_transmit_bytes_total", "Bytes sent since boot.", float64(network.Usage.TxBytes), labels...)
}

if tcp := data.TCPStates; tcp != nil {
states := []struct {
name string
count int
}{
{"established", tcp.Established}, {"syn_sent", tcp.SynSent}, {"syn_recv", tcp.SynRecv},
{"fin_wait1", tcp.FinWait1}, {"fin_wait2", tcp.FinWait2}, {"time_wait", tcp.TimeWait},
{"close", tcp.Close}, {"close_wait", tcp.CloseWait}, {"last_ack", tcp.LastAck},
{"listen", tcp.Listen}, {"closing", tcp.Closing},
}
for _, state := range states {
m.gauge("symon_tcp_connections", "TCP connections by state.", float64(state.count), "host", host, "state", state.name)
}
}

if p := data.Pressure; p != nil {
resources := []struct {
name string
pressure monitor.ResourcePressure
}{{"cpu", p.CPU}, {"memory", p.Memory}, {"io", p.IO}}
for _, r := range resources {
m.gauge("symon_pressure_some_percent", "Share of the last 10 seconds some tasks waited on the resource.", r.pressure.Some.Avg10, "host", host, "resource", r.name)
if r.pressure.FullAvailable {
m.gauge("symon_pressure_full_percent", "Share of the last 10 seconds all tasks waited on the resource.", r.pressure.Full.Avg10, "host", host, "resource", r.name)
}
}
}

for _, temp := range data.Temperatures {
sensor := temp.Name
if temp.Label != "" {
sensor += "/" + temp.Label
}
m.gauge("symon_temperature_celsius", "Sensor temperature.", temp.Celsius, "host", host, "sensor", sensor)
}
for _, service := range data.Services {
m.gauge("symon_service_up", "1 when the service is running.", boolValue(service.Running), "host", host, "service", service.Name)
}

for _, container := range data.Containers {
name := container.Name
if name == "" {
name = container.ShortID
}
labels := []string{"host", host, "container", name, "project", container.ComposeProject}
m.gauge("symon_container_cpu_percent", "Container CPU usage, as a share of the host.", container.CPU.PercentOfHost, labels...)
m.gauge("symon_container_memory_bytes", "Container memory in use.", container.Memory.Used, labels...)
if rates := container.Rates; rates != nil {
// no traffic rates for containers on the host network
if rates.RxBytesPerSec != nil {
m.gauge("symon_container_receive_bytes_per_second", "Container network traffic received.", *rates.RxBytesPerSec, labels...)
}
if rates.TxBytesPerSec != nil {
m.gauge("symon_container_transmit_bytes_per_second", "Container network traffic sent.", *rates.TxBytesPerSec, labels...)
}
m.gauge("symon_container_read_bytes_per_second", "Container disk reads.", rates.ReadBytesPerSec, labels...)
m.gauge("symon_container_write_bytes_per_second", "Container disk writes.", rates.WriteBytesPerSec, labels...)
}
}
}

// parsePercent turns "40%" into 40
func parsePercent(value string) (float64, bool) {
pct, err := strconv.ParseFloat(strings.TrimSuffix(strings.TrimSpace(value), "%"), 64)
return pct, err == nil
}

func boolValue(b bool) float64 {
if b {
return 1
}
return 0
}

// metricsWriter builds the Prometheus text format. Each metric's HELP,
// TYPE and samples have to be written together, so samples are grouped by
// metric and written at the end.
type metricsWriter struct {
families []*metricFamily
byName map[string]*metricFamily
}

type metricFamily struct {
name string
kind string
help string
samples []string
// labels already written, a repeated set would make the output invalid
seen map[string]bool
}

func newMetricsWriter() *metricsWriter {
return &metricsWriter{byName: map[string]*metricFamily{}}
}

func (w *metricsWriter) gauge(name string, help string, value float64, labels ...string) {
w.add(name, "gauge", help, value, labels)
}

func (w *metricsWriter) counter(name string, help string, value float64, labels ...string) {
w.add(name, "counter", help, value, labels)
}

var labelEscaper = strings.NewReplacer(`\`, `\\`, `"`, `\"`, "\n", `\n`)

// add records one sample. labels are name and value pairs.
func (w *metricsWriter) add(name string, kind string, help string, value float64, labels []string) {
family, ok := w.byName[name]
if !ok {
family = &metricFamily{name: name, kind: kind, help: help, seen: map[string]bool{}}
w.byName[name] = family
w.families = append(w.families, family)
}

var set strings.Builder
for i := 0; i+1 < len(labels); i += 2 {
if i > 0 {
set.WriteByte(',')
}
fmt.Fprintf(&set, `%s="%s"`, labels[i], labelEscaper.Replace(labels[i+1]))
}
if family.seen[set.String()] {
return
}
family.seen[set.String()] = true

sample := name
if set.Len() > 0 {
sample += "{" + set.String() + "}"
}
family.samples = append(family.samples, sample+" "+strconv.FormatFloat(value, 'g', -1, 64)+"\n")
}

func (w *metricsWriter) writeTo(out io.Writer) error {
var b strings.Builder
for _, family := range w.families {
fmt.Fprintf(&b, "# HELP %s %s\n# TYPE %s %s\n", family.name, family.help, family.name, family.kind)
for _, sample := range family.samples {
b.WriteString(sample)
}
}
_, err := io.WriteString(out, b.String())
return err
}
92 changes: 92 additions & 0 deletions client/internal/server/metrics_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
package server

import (
"context"
"encoding/json"
"net/http/httptest"
"strings"
"testing"

"github.com/dhamith93/SyMon/internal/api"
"github.com/dhamith93/SyMon/internal/monitor"
)

// Snapshots has web1 reporting, db1 gone quiet and new1 not sent anything yet
func (f *fakeCollector) Snapshots(ctx context.Context, in *api.Void) (*api.SnapshotList, error) {
snapshot, err := json.Marshal(monitor.MonitorData{
System: monitor.System{UpTimeSeconds: 3600},
ProcUsage: monitor.CPU{LoadAvg: 37, Load1: 0.5},
Memory: monitor.Memory{PercentageUsed: 42.5, Total: 16000, Available: 9000},
Disk: []monitor.Disk{
{FileSystem: "/dev/sda1", MountedOn: "/", Usage: monitor.DiskUsage{Size: 1000, Used: 400, Usage: "40%"}, Inodes: monitor.InodeUsage{Usage: "10%"}},
},
Networks: []monitor.Network{{Interface: "eth0", Usage: monitor.NetworkUsage{RxBytes: 1000, TxBytes: 2000}}},
Services: []monitor.Service{{Name: "nginx", Running: true}},
Containers: []monitor.Container{
{ShortID: "aaa111", Name: "web", ComposeProject: "shop", CPU: monitor.ContainerCPU{PercentOfHost: 12},
Rates: &monitor.ContainerRates{RxBytesPerSec: floatPtr(2048), ReadBytesPerSec: 10}},
},
})
if err != nil {
return nil, err
}
return &api.SnapshotList{Hosts: []*api.HostSnapshot{
{Host: "web1", Up: true, LastSeen: 1700000000, Time: 1700000000, SnapshotJson: string(snapshot)},
{Host: "db1", LastSeen: 1690000000, Time: 1690000000, SnapshotJson: string(snapshot)},
{Host: "new1", Up: true, LastSeen: 1700000000},
}}, nil
}

func TestMetrics(t *testing.T) {
s, _ := newTestServer(t, nil)
rec := httptest.NewRecorder()
s.routes().ServeHTTP(rec, httptest.NewRequest("GET", "/metrics", nil))
body := rec.Body.String()

if rec.Code != 200 || rec.Header().Get("Content-Type") != "text/plain; version=0.0.4; charset=utf-8" {
t.Fatalf("unexpected response %d %q: %s", rec.Code, rec.Header().Get("Content-Type"), body)
}
for _, line := range []string{
"# TYPE symon_up gauge\n" + `symon_up{host="web1"} 1` + "\n" + `symon_up{host="db1"} 0` + "\n" + `symon_up{host="new1"} 1`,
`symon_last_seen_timestamp_seconds{host="db1"} 1.69e+09`,
`symon_cpu_usage_percent{host="web1"} 37`,
`symon_memory_total_bytes{host="web1"} 1.6777216e+10`,
`symon_disk_used_percent{host="web1",device="/dev/sda1",mount="/"} 40`,
"# TYPE symon_network_receive_bytes_total counter\n" + `symon_network_receive_bytes_total{host="web1",iface="eth0"} 1000`,
`symon_service_up{host="web1",service="nginx"} 1`,
`symon_container_receive_bytes_per_second{host="web1",container="web",project="shop"} 2048`,
} {
if !strings.Contains(body, line+"\n") {
t.Errorf("expected %q in:\n%s", line, body)
}
}
// db1 stopped reporting, so its old values are left out
if strings.Contains(body, `{host="db1",`) || strings.Contains(body, `symon_cpu_usage_percent{host="db1"}`) {
t.Errorf("expected only up and last seen for db1:\n%s", body)
}
// web1 has no transmit rate, as if it were on the host network
if strings.Contains(body, "symon_container_transmit_bytes_per_second") {
t.Errorf("expected no transmit rate:\n%s", body)
}
}

func TestMetricsWriter(t *testing.T) {
w := newMetricsWriter()
w.gauge("a", "First.", 1, "name", `quote " backslash \ newline`+"\n")
w.counter("b_total", "Second.", 2)
w.gauge("a", "First.", 3, "name", "other")
// a repeated label set is dropped, Prometheus rejects the whole scrape otherwise
w.gauge("a", "First.", 4, "name", "other")

var out strings.Builder
if err := w.writeTo(&out); err != nil {
t.Fatal(err)
}
want := "# HELP a First.\n# TYPE a gauge\n" +
`a{name="quote \" backslash \\ newline\n"} 1` + "\n" +
`a{name="other"} 3` + "\n" +
"# HELP b_total Second.\n# TYPE b_total counter\nb_total 2\n"
if out.String() != want {
t.Errorf("got:\n%s\nwant:\n%s", out.String(), want)
}
}
Loading
Loading