Smart Messenger for Auto-waking Unwake GHz
Config-driven reverse proxy with automatic Wake-on-LAN for homelab services
"I am fire. I am... asleep until you need me."
Smaug is a lightweight, production-ready reverse proxy designed for homelab environments. It routes requests to backend services transparently while automatically waking sleeping servers via Wake-on-LAN and putting idle servers to sleep reducing power consumption by 90%+ during idle periods.
Typical homelab GPU servers (Ollama, Marker, Whisper) consume 200-400W when idle. Smaug lets them sleep when not in use and wakes them on-demand, enabling efficient resource management without sacrificing responsiveness.
- Transparent Proxying: Routes requests to backend services via port-per-service model
- Auto-Wake: Detects sleeping servers and sends WoL packets automatically
- Auto-Sleep: Puts idle servers to sleep after configurable timeout
- YAML Configuration: Simple, file-driven setup with hot-reload support
- Observability: Prometheus metrics, structured JSON logging, health checks
- Privilege Separation: Unprivileged proxy + privileged Gwaihir WoL service
- Kubernetes Native: Built for K8s deployments with minimal resource footprint
- Production Ready: Graceful shutdown, rate limiting, security best practices
graph TB
subgraph K8s["K8s Cluster (ai namespace)"]
subgraph SMAUG_Box["SMAUG (Unprivileged)"]
ConfigMgr["Config Manager<br/>(hot-reload)"]
RouteMgr["Route Manager<br/>(per-port listen)"]
ProxyHandler["Proxy Handler"]
HealthChecker["Health Checker"]
IdleTracker["Idle Tracker"]
MetricsExp["Metrics Exporter<br/>(:2112)"]
end
subgraph Gwaihir_Box["Gwaihir (Privileged)"]
WoLSender["WoL Sender"]
end
Clients["Clients<br/>(OpenWebUI, Agents)"]
end
Backend["Backend Server<br/>(Ollama, Marker, etc)<br/>GPU Workload<br/>WoL-Enabled"]
Clients -->|HTTP Request| SMAUG_Box
ConfigMgr --> RouteMgr
RouteMgr --> ProxyHandler
ProxyHandler --> HealthChecker
HealthChecker --> IdleTracker
IdleTracker --> MetricsExp
HealthChecker -->|Check Health| Backend
IdleTracker -->|API Call<br/>POST /wol| Gwaihir_Box
Gwaihir_Box -->|UDP :9<br/>WoL Magic Packet| Backend
IdleTracker -->|Sleep Request| Backend
ProxyHandler -->|HTTP Proxy| Backend
Backend -->|Response| Clients
style K8s fill:#e1f5ff
style SMAUG_Box fill:#fff3e0
style Gwaihir_Box fill:#f3e5f5
style Backend fill:#e8f5e9
When a client requests a backend service, SMAUG orchestrates the wake-up:
sequenceDiagram
participant Client
participant SMAUG
participant HealthCheck as Health<br/>Checker
participant Gwaihir
participant Backend
Client->>SMAUG: Request (port 11434)
SMAUG->>HealthCheck: Check backend health
alt Backend Unhealthy
HealthCheck->>Gwaihir: POST /wol<br/>(Machine ID: backend)
Gwaihir->>Backend: UDP :9<br/>WoL Magic Packet
HealthCheck->>Backend: Poll health<br/>(2s interval)
Note over HealthCheck,Backend: Retry until healthy<br/>or timeout (60s)
Backend-->>HealthCheck: 200 OK /api/tags
else Backend Already Healthy
Backend-->>HealthCheck: 200 OK /api/tags
end
SMAUG->>Backend: Proxy request<br/>HTTP GET /api/chat
Backend-->>SMAUG: Response (AI output)
SMAUG-->>Client: Transparent proxy response
Note over SMAUG: Track request for<br/>idle detection
Typical wake latency: 30-90 seconds (depends on server hardware)
An idle tracker monitors request activity and puts idle servers to sleep:
sequenceDiagram
participant IdleTracker
participant SMAUG
participant Backend
Note over IdleTracker: Monitor request activity
IdleTracker->>IdleTracker: No traffic for 5m?
alt Idle Timeout Reached
IdleTracker->>Backend: POST /sleep<br/>(graceful shutdown)
Backend->>Backend: Save state
Backend->>Backend: Close connections
Backend-->>IdleTracker: 200 OK
Note over Backend: Server powers down
IdleTracker->>SMAUG: Emit metric<br/>sleep_triggered
else Recent Traffic
Note over IdleTracker: Reset idle timer<br/>Continue monitoring
end
# Global settings
settings:
gwaihir:
url: "http://gwaihir-service.ai.svc.cluster.local"
apiKey: "${GWAIHIR_API_KEY}" # Env var substitution
timeout: 5s
logging:
level: info
format: json # structured logging
# Observability configuration
# Controls which infrastructure endpoints are exposed
observability:
# Health check endpoints: /health, /live, /ready, /version
healthCheck:
enabled: true
port: 2111
# Metrics endpoint: /metrics (Prometheus format)
metrics:
enabled: true
port: 2112
# Server definitions
servers:
saruman:
wakeOnLan:
enabled: true
machineId: "saruman" # Gwaihir machine ID
timeout: 60s
debounce: 5s # Min interval between WoL attempts
sleepOnLan:
enabled: true
endpoint: "http://saruman.from-gondor.com:8000/sleep"
authToken: "${SLEEP_ON_LAN_TOKEN}" # Env var substitution
idleTimeout: 5m # Sleep after 5 min idle
healthCheck:
endpoint: "http://saruman.from-gondor.com:8000/status"
authToken: "${HEALTH_CHECK_TOKEN}" # Optional: base64-encoded user:password for Basic Auth
interval: 2s
timeout: 2s
# Route definitions (port-per-service)
routes:
- name: ollama
listen: 11434
upstream: "http://saruman.from-gondor.com:11434"
server: saruman
- name: marker
listen: 8080
upstream: "http://saruman.from-gondor.com:8080"
server: sarumanSMAUG_CONFIG- Path to configuration file (default:/etc/smaug/services.yaml)GWAIHIR_API_KEY- API key for Gwaihir service (required if WoL enabled)SMAUG_LOG_LEVEL- Log level: debug|info|warn|error (default: info)SMAUG_LOG_FORMAT- Log format: json|text (default: json)
- Go 1.26+ (for building)
- Backend servers with Wake-on-LAN enabled
- Gwaihir service deployed (for WoL functionality)
- Network L2 adjacency for WoL broadcast (same subnet/VLAN)
# Clone repository
git clone https://github.com/josimar-silva/smaug.git
cd smaug
# Install dependencies
go mod download
# Run tests
just test
# Build binary
just build
# Run locally
export SMAUG_CONFIG=config/smaug.example.yaml
export GWAIHIR_API_KEY=your-secret-key
just rundocker run -d \
--name smaug \
--network host \
-v /path/to/config.yaml:/etc/smaug/services.yaml:ro \
-e GWAIHIR_API_KEY=your-secret-key \
-e SMAUG_LOG_LEVEL=info \
ghcr.io/josimar-silva/smaug:latestSee docs/adrs/ADR-001-foundation.md for complete Kubernetes deployment specifications including Deployment, Service, ConfigMap, NetworkPolicy, and RBAC resources.
Quick start:
# Create ConfigMap with services.yaml
kubectl create configmap smaug-config --from-file=services.yaml
# Apply manifests
kubectl apply -f deploy/
# Verify
kubectl logs -f deployment/smaug
kubectl port-forward svc/smaug 2112:2112
curl http://localhost:2112/metricsinternal/
├── proxy/ # Proxy and routes
├── store/ # Storage
├── health/ # Server health checker
├── config/ # Configuration parsing
├── middleware/ # HTTP middleware
├── infrastructure/ # Logging, metrics, etc.
├── management/ # Server Management
└── client/ # External service clients (Gwaihir, sleep)
cmd/smaug/ # Application entry point
tests/ # Integration tests
config/ # Example configurations
# Format and lint
just format
just lint
# Run all checks
just pre-commit
# Run tests with coverage
just test
# Run with race detector
go test -race ./...
# Build binary
just build
# Build Docker image
just docker-build latestKey decisions documented in ADRs:
Available on port 2112 (/metrics endpoint).
Key metrics:
smaug_wake_attempts_total- WoL wake attempts by server and success statussmaug_server_awake- Server state gauge (1=awake, 0=sleeping) by serversmaug_health_check_failures_total- Health check failures by server and reasonsmaug_sleep_triggered_total- Servers put to sleep due to idle timeoutsmaug_request_duration_seconds- Proxy request latency (histogram)smaug_requests_total- Total HTTP requests processedsmaug_requests_by_status_total- HTTP requests by status codesmaug_requests_by_method_total- HTTP requests by methodsmaug_config_reload_total- Config reload attempts by success statussmaug_gwaihir_api_calls_total- Gwaihir API calls by operation and success statussmaug_gwaihir_api_duration_seconds- Gwaihir API call duration (histogram)
JSON logs for log aggregation (ELK, Loki, etc).
Example log entry:
{
"level": "info",
"timestamp": "2026-02-11T10:30:45Z",
"msg": "wol_sent",
"server": "saruman",
"duration_ms": 120
}Management server runs on port 2111 (configurable) and exposes:
/live: Kubernetes liveness probe — returns{"status":"alive"}when the process is running/ready: Kubernetes readiness probe — returns 503 when no routes are active/health: Overall application health including active route count and uptime/version: Build version, git commit, and build time
K8s probes configured in deployment manifest.
- Privilege Separation: SMAUG runs unprivileged, Gwaihir handles WoL with elevated privileges
- Network Isolation: NetworkPolicy restricts SMAUG → Gwaihir, Prometheus → metrics
- Authentication: Gwaihir API key stored in K8s Secret (not ConfigMap)
- Input Validation: Config schema validation, URL parsing
- Rate Limiting: 1 WoL request per 10s per server, prevents DoS attacks
- Secret Redaction: API keys and auth tokens are redacted in logs and config marshalling
| Threat | Mitigation |
|---|---|
| Gwaihir API abuse | API key auth + NetworkPolicy + rate limiting |
| Config tampering | RBAC + validation + hash verification |
| Wake flooding DoS | Rate limiting (1/10s per server) + debounce |
| Sleep API SSRF | Endpoint validation + scheme validation |
| Metrics info disclosure | NetworkPolicy (Prometheus only) |
| Secret leakage in logs | SecretString type redacts values in all output |
See ADR-001: Security Considerations for full threat model.
| Metric | Idle | Active (10 req/s) |
|---|---|---|
| CPU | 10m | 100m |
| Memory | 50Mi | 128Mi |
| Network | <1 Mbps | Depends on backend responses |
| Scenario | Latency |
|---|---|
| Backend awake, healthy | ~5ms (proxy overhead) |
| Backend sleeping, wake succeeds | 30-90s (WoL + boot time) |
| Backend unreachable | 2s (health check timeout) |
- Routes: No practical limit (each route is independent listener)
- Servers: Tested with 10+ servers without issues
- Concurrency: Limited by backend capacity (SMAUG adds minimal overhead)
See ADR-001: Performance Characteristics for detailed benchmarks.
We welcome contributions! See CONTRIBUTING.md for guidelines on:
- Code standards (SOLID, functional principles)
- Test-driven development (TDD)
- Commit practices (atomic commits, conventional commits)
- Pull request workflow
- Go 1.26+
justfor task runninggolangci-lintfor linting- Docker for container builds
Before submitting PRs:
# Format code
just format
# Run linters
just lint
# Run all pre-commit checks
just pre-commit
# Ensure 90%+ test coverage
just test- Verify Gwaihir is deployed and reachable
- Check
GWAIHIR_API_KEYis set correctly - Verify backend has WoL enabled in BIOS
- Ensure network L2 adjacency (same subnet/VLAN)
- Check logs:
kubectl logs -f deployment/smaug | grep wol
- Verify
sleepOnLan.enabled: truein config - Check backend has sleep endpoint (e.g.,
POST /sleep) - Verify
idleTimeoutis reasonable (default 5m) - Monitor logs for
sleep_triggeredevents
- Increase
healthCheck.timeoutif backend is slow to respond - Increase
wakeOnLan.timeoutif servers take > 60s to wake - Check backend logs for errors
- Monitor latency with
smaug_request_duration_secondsmetric
- Validate YAML syntax:
yamllint services.yaml - Check logs for
config_reload_failederrors - Previous config is kept on reload failure (no downtime)
- Review Configuration section
- Gwaihir - Privileged WoL messenger service
- ADRs - Comprehensive design decisions, deployment specs, security model, observability strategy
- CONTRIBUTING.md - Contribution process and standards
This project is licensed under the MIT License - see LICENSE file for details.
