Reference for the two extension surfaces: the module contract (Go interface implemented by host capabilities) and the manager API (REST/JSON consumed by clients).
| Boundary | Protocol |
|---|---|
| manager ↔ associate | persistent gRPC bidirectional stream over mTLS (associate-initiated) |
| manager ↔ clients (dashboard, CLI, Home Assistant, webhooks) | REST/JSON over TLS |
| associate ↔ modules | in-process Go interface (compiled-in) |
A module is a Go package implementing Module. All metadata required by the manager and dashboard
is exposed through Manifest(); no core changes are needed to add a module.
type Module interface {
Manifest() Manifest
Detect(ctx context.Context) (Detection, error)
Status(ctx context.Context) (Status, error)
Execute(ctx context.Context, req ActionRequest, emit func(Event)) (Result, error)
}| Method | Mutating | Description |
|---|---|---|
Manifest |
no | Static metadata and the list of supported actions. |
Detect |
no | Reports applicability to the host and detected capabilities. |
Status |
no | Read-only/dry-run snapshot of current state. |
Execute |
yes | Runs a named action; emit streams progress/log events. |
type Manifest struct {
Name string
Version string
Description string
Actions []ActionSpec
ConfigSchema json.RawMessage // JSON Schema for module-level per-host config (optional)
}
type ActionSpec struct {
Name string
Description string
ParamsSchema json.RawMessage // JSON Schema for request params
ResultSchema json.RawMessage // JSON Schema for the result payload
Privilege Privilege // None | Elevated
Destructive bool
DefaultTimeout time.Duration
Streams bool
}
type Privilege int // None, Elevated
type Detection struct {
Applicable bool
Capabilities map[string]string // e.g. {"distro": "debian", "orchestrator": "compose"}
}
type Status struct {
Summary string
Data json.RawMessage
}
type ActionRequest struct {
JobID string
Action string
Params json.RawMessage
}
type Event struct {
JobID string
Kind EventKind // Log | Progress | State
Message string
Progress float64 // 0.0–1.0 when Kind == Progress
}
type Result struct {
State JobState // Succeeded | Failed | TimedOut
Data json.RawMessage
}| Field | Type | Notes |
|---|---|---|
Name |
string | Unique within the module. Used in the action endpoint path. |
ParamsSchema |
JSON Schema | Validated before dispatch; drives generic dashboard form rendering. |
ResultSchema |
JSON Schema | Shape of Result.Data. |
Privilege |
enum | Elevated actions run via the associate's privileged helper. |
Destructive |
bool | Requires confirmation to run manually; requires explicit opt-in to schedule. |
DefaultTimeout |
duration | Overridable per task. |
Streams |
bool | Whether the action emits Events during execution. |
- Privilege is declared per action, not per module. Modules request elevation; the associate owns the single privileged helper.
- Each
Executecarries a manager-issuedJobID. The associate serializes actions per host and rejects an identical action already queued or running (idempotency). Destructiveactions require confirmation when run manually and explicit opt-in to be scheduled.- A module may declare
Manifest.ConfigSchemafor per-host configuration (e.g. a private-registry credential forduo). The manager stores this config per host and serves it via the module config endpoints; config secrets are kept outsidestate.json. - Built-in modules:
duo,qup,sys. - External modules:
Manifest/ActionSpecmap to a JSON/stdio (or local gRPC) protocol; an out-of-process module loader can be added without changing the interface.
Base path: /api/v1. All requests and responses are JSON over TLS.
Every endpoint requires authentication via one of:
- Session cookie established by
POST /auth/login. Authorization: Bearer <token>— tokens are minted and revoked via/auth/tokens.
- List endpoints paginate with
?limit=and?cursor=; responses includenext_cursor. - Server→client streams use Server-Sent Events (
Accept: text/event-stream). - Timestamps are RFC 3339 UTC.
Non-2xx responses use:
{ "error": { "code": "string", "message": "string" } }| Status | Meaning |
|---|---|
| 400 | Invalid request / schema validation failure |
| 401 | Missing or invalid authentication |
| 403 | Authenticated but not permitted |
| 404 | Resource not found |
| 409 | Conflict (e.g. duplicate in-flight action) |
| 422 | Action rejected by policy (e.g. unconfirmed destructive action) |
| Method & path | Description |
|---|---|
POST /auth/login |
Establish a session. |
POST /auth/logout |
End the session. |
POST /auth/tokens |
Mint an API token. |
DELETE /auth/tokens/{id} |
Revoke an API token. |
GET /overview |
Lab-wide aggregate summary (see Overview). |
GET /hosts |
List hosts. |
POST /hosts |
Start async host enrollment (see Host lifecycle). |
GET /hosts/{id} |
Get host detail. |
PUT /hosts/{id} |
Update editable host fields (ip, tailscale, ssh_user, enabled modules). |
DELETE /hosts/{id} |
Remove a host (revokes its cert). |
POST /hosts/{id}/upgrade |
Push the manager's current associate build to the host and restart it (keeps cert/bundle). Returns { "jobId": "..." }. |
POST /hosts/upgrade-stale |
Same, for every host running older associate code — hosts whose associate code already matches are skipped, even when the manager is newer. Takes no body: each host is tried with the manager's SSH key/agent and its own stored sshUser. Returns { "started": [{ "hostId", "hostName", "jobId" }], "failed": [{ "hostId", "hostName", "error" }] } — one job per host, failures reported per host rather than aborting the batch. A host whose job fails for want of credentials has { "authRequired": true } in its job result; retry that host through POST /hosts/{id}/upgrade with a login it accepts. The flag is a hint, not a filter — it is derived from the host's error text, so any failed host is worth offering a retry. |
POST /hosts/refresh |
Ask every connected associate for a heartbeat and all module statuses right now instead of waiting for the next tick. Returns 202 { "requested": n } (hosts the request reached); the updates arrive on the event stream as host_updated. |
GET /hosts/{id}/status |
Live health and per-module states. |
GET /hosts/{id}/modules |
Enabled modules with manifests and detection results. |
POST /hosts/{id}/modules/{name}:enable |
Enable a module on the host. |
POST /hosts/{id}/modules/{name}:disable |
Disable a module on the host. |
GET /hosts/{id}/modules/{name}/config |
Read a module's per-host config + schema. |
PUT /hosts/{id}/modules/{name}/config |
Update a module's per-host config (validated against schema). |
POST /hosts/{id}/modules/{name}/actions/{action} |
Start an action. Returns { "jobId": "..." }. |
GET /services |
Aggregate of compose stacks + services across all hosts (see Services). |
GET /jobs |
List jobs. Filter with ?host=, ?status=. |
GET /jobs/{id} |
Get job detail and result. |
GET /jobs/{id}/events |
SSE stream of progress and log events. |
GET /approvals |
List pending approvals. |
POST /approvals/{id}:confirm |
Confirm a pending action. |
POST /approvals/{id}:reject |
Reject a pending action. |
GET /tasks |
List scheduled tasks. |
POST /tasks |
Create a scheduled task. |
PUT /tasks/{id} |
Update a scheduled task. |
DELETE /tasks/{id} |
Delete a scheduled task. |
POST /tasks/{id}/run |
Run a task now, outside its schedule (even if disabled). Next run is unchanged. |
GET /hosts/{id}/files?path= |
Read a config file. |
PUT /hosts/{id}/files |
Validate, back up, and write a config file. |
POST /hosts/{id}/files:undo |
Restore the last local copy. |
GET /hosts/{id}/logs |
SSE stream of host or container logs. |
GET /audit |
Read the audit log. |
GET /events |
SSE stream of host online/offline, job, and status updates. |
GET /settings |
Read manager settings. |
PUT /settings |
Update manager settings. Omitted fields keep their current value. heartbeatSeconds (default 30) and statusPollSeconds (default 60) set how often associates send vitals and re-poll module statuses; both must be 5–3600 and are pushed to connected associates immediately. |
GET /manager/version |
Compare the running manager with the tip of its checkout's branch on origin (read-only git ls-remote, cached ~10 min; ?refresh=1 forces a fresh check). Returns { "running", "local", "remote", "branch", "updateAvailable", "checkedAt", "error"? }; error explains why no comparison was possible (not a checkout, detached HEAD, remote unreachable). |
GET /backup |
Export settings. |
POST /restore |
Import settings. |
POSTto an action endpoint creates a job in statequeuedand returns itsjobId.- If the action's
Destructiveflag is set, or policy requires approval, the manager creates a pending approval and holds the job; otherwise it dispatches immediately. POST /approvals/{id}:confirmdispatches the command over the associate's mTLS stream.- The associate executes the action; clients follow
GET /jobs/{id}/eventsfor progress and logs. - The manager records the
Resultand writes an audit entry.
queued → running → (succeeded | failed | timed_out)
GET /events is the aggregate live feed; clients subscribe to it for host, job, and status changes
rather than polling. GET /jobs/{id}/events and GET /hosts/{id}/logs are scoped streams.
A host's state is one of enrolling | online | offline | error.
Enrollment is asynchronous. POST /hosts accepts:
{ "ip": "10.0.0.5", "tailscale": true, "ssh_user": "admin", "ssh_password": "…", "modules": ["duo", "qup", "sys"] }ssh_password is optional and transient — used only for the SSH bootstrap, never persisted to
state.json. The call returns immediately with the host in enrolling state plus a jobId:
{ "host": { "id": "h1", "state": "enrolling" }, "jobId": "…" }Flow: POST /hosts → SSH → install associate → mTLS exchange → associate dials home → online.
Progress is reported on GET /jobs/{id}/events and the aggregate GET /events. Failures move the
host to error with detail in the job result.
GET /overview returns a lab-wide aggregate for the dashboard's default page (shape may be refined):
{
"hosts": { "total": 8, "online": 7, "offline": 1, "enrolling": 0 },
"updates": { "packages": 23, "images": 4 },
"resources": { "cpu_percent": 31.2, "mem_percent": 44.0 },
"services": { "total": 40, "running": 38, "stopped": 2 }
}GET /services is a read-only projection over the duo module across all hosts — compose stacks
with their nested services:
{
"stacks": [
{
"host_id": "h1", "name": "media", "path": "/srv/media/compose.yaml", "status": "running",
"services": [
{ "name": "jellyfin", "status": "running", "image": "jellyfin/jellyfin:latest", "has_logs": true },
{ "name": "sonarr", "status": "stopped", "has_logs": true }
]
}
]
}Control is not a separate surface — start/stop/restart route through duo actions on the owning
host:
POST /hosts/{host_id}/modules/duo/actions/{start|stop|restart}
{ "stack": "media", "service": "jellyfin" }Omit service for stack-level control; include it for a single service (both levels are supported).
Logs use GET /hosts/{host_id}/logs with service/container parameters. (v1 covers compose services
only; a user-defined non-docker service registry is deferred.)
The dashboard's compose editor reads and writes a stack's files through duo actions. The compose
file is the one recorded on the stack's containers (com.docker.compose.project.config_files); the
.env is the project env file next to it, which docker compose reads for ${VAR} substitution.
| Action | Params | Result (job.result) |
Notes |
|---|---|---|---|
read-compose |
stack |
stack, path, content, truncated, multiFile, sha256 |
Read-only. Content is capped at 1 MiB (truncated). |
write-compose |
stack, content, baseSha256? |
sha256 |
Keeps <file>.bak, validates with docker compose config, restores on failure. Refused for multi-file stacks. |
read-env |
stack |
stack, path, target, outsideDir, content, exists, truncated, sha256 |
Read-only. A symlinked .env is followed: target is the real file, outsideDir flags one outside the stack directory. Missing file: exists: false. |
write-env |
stack, content, baseSha256? |
sha256, created, target |
Validates the candidate first (docker compose --env-file <candidate> config), so a rejected edit never touches the live file. Keeps <target>.bak; creates a missing file with mode 0600. Capped at 256 KiB. |
deploy |
stack, service?, removeOrphans? |
— | docker compose -p <stack> -f <file>… up -d, with every compose file the stack was created from, so it always acts on the stack's own project; removeOrphans adds --remove-orphans. Destructive: queued for approval. |
baseSha256is thesha256from the last read or write. When set, the write fails if the file changed on the host in the meantime (including being created or deleted) instead of overwriting it.write-envis deliberately not destructive: approvals record their params in the audit log, and these params are the file's contents. Its validation errors have.envvalues masked.- Hosts whose associate predates
read-env/write-envdon't list them in their module manifest; clients should check the manifest rather than dispatch and fail. Likewise, an associate whosedeployparams schema has noremoveOrphansignores that param.
GET /hosts/{id}/modules/{name}/config returns the stored config and its schema (from
Manifest.ConfigSchema), so a client can render a generic settings form:
{ "config": { "…": "…" }, "schema": { "…": "JSON Schema" } }PUT validates the body against the schema before storing. Config secrets are persisted outside
state.json.
GET /settings / PUT /settings manage manager-wide configuration (listen/TLS, scheduler default
timezone, audit retention, etc.). User credentials and API tokens are managed via the /auth/*
endpoints; the audit log and backup/restore have their own endpoints above.