A portfolio-grade, local-first observability control plane that turns service metrics, alerts, distributed traces, structured logs, and SLO error budgets into one explainable incident story.
Live demo: cloud-ops-observability-lab.onrender.com — the free instance may need up to 50 seconds to wake after inactivity.
The project demonstrates modern platform engineering without requiring a cloud account or touching production infrastructure. Four deterministic scenarios let reviewers explore healthy operations, latency, authentication failures, and a dependency outage safely.
Modern IT operations is moving beyond isolated dashboards toward correlated, machine-readable telemetry that humans and agents can reason about. This lab demonstrates that workflow while keeping the important operational boundary explicit: observation and explanation are allowed; infrastructure changes are not.
It is designed to show employers that the author can connect software engineering and IT operations:
- instrument a real Node.js service with the OpenTelemetry SDK;
- model metrics, traces, logs, alerts, incidents, timelines, and SLOs with TypeScript;
- correlate evidence across services instead of presenting unrelated charts;
- design structured JSON APIs for automation and AI-assisted operations;
- build an accessible responsive React dashboard;
- ship tests, CI, containers, deployment metadata, security notes, and documentation.
- Five-service topology with owners, versions, regions, dependencies, health states, and golden signals.
- Four repeatable scenarios: healthy baseline, latency regression, authentication errors, and dependency outage.
- Cross-signal incident correlation with severity, likely cause, confidence, evidence, affected services, trace links, alert links, and a timeline.
- Distributed trace waterfall with parent-child spans, attributes, duration, and error status.
- Structured log explorer with level filters, text search, trace IDs, span IDs, and attributes.
- SLO and error-budget view with a 30-day window, remaining budget, burn rate, and policy state.
- Real runtime telemetry captured by the official OpenTelemetry Node SDK for Express requests.
- Local runbook library exposed through read-only endpoints.
- JSON export for downstream automation or incident-analysis tools.
- Safe by default: no credentials, cloud calls, shell execution, write actions, or remediation endpoints.
| Scenario | Customer impact | Signals to inspect | Correlated outcome |
|---|---|---|---|
| Healthy baseline | None | Stable golden signals and budgets | No incident or alert |
| Latency regression | Slow checkout | Inventory latency, checkout burn, warning logs, slow database span | SEV-2 with 91% confidence |
| Authentication errors | Some sign-ins fail | Identity error rate, edge 401s, signing-key errors | SEV-2 with 88% confidence |
| Dependency outage | Checkout and notifications disrupted | Availability collapse, retries, failed spans, exhausted budgets | SEV-1 with 96% confidence |
flowchart LR
U[React dashboard] -->|GET scenario snapshot| A[Express API]
A --> V[Zod validation]
V --> S[Deterministic scenario engine]
S --> M[Metrics and SLO model]
S --> C[Incident correlation]
S --> T[Simulated distributed traces]
S --> L[Structured scenario logs]
A --> R[Local runbook JSON]
A --> O[OpenTelemetry Node SDK]
O --> E[In-memory span exporter]
E --> X[Runtime telemetry view]
M --> J[Typed observability snapshot]
C --> J
T --> J
L --> J
R --> J
X --> J
J --> U
style U fill:#10233a,stroke:#58a6ff,color:#fff
style J fill:#12352d,stroke:#42d392,color:#fff
style O fill:#2d2552,stroke:#9d8cff,color:#fff
The dashboard performs one snapshot request per refresh. Every panel renders from that same typed response, preventing the metrics, incident narrative, trace, logs, and SLO state from contradicting one another.
| Area | Implementation |
|---|---|
| Express request spans | Real OpenTelemetry SDK spans with non-zero trace and span IDs |
| Request logs | Real structured JSON runtime logs in a bounded memory buffer |
| Service metrics and SLO history | Deterministic local scenario data |
| Distributed service traces | Deterministic cross-service examples for investigation practice |
| Runbooks | Bundled local JSON; never fetched or executed remotely |
| Infrastructure actions | Intentionally absent |
| Technology | Why it was chosen |
|---|---|
| TypeScript | One language and shared types across the server, simulator, tests, and UI |
| Express 5 | Small, readable API surface with familiar HTTP concepts |
| React 19 + Vite | Fast local feedback and component-based dashboard development |
| OpenTelemetry JS | Vendor-neutral observability concepts and official runtime instrumentation |
| Zod | Clear validation at the API boundary |
| Node test runner | Reliable tests without a large additional test framework |
You do not need prior experience with these tools to run the project. The first useful files to read are src/scenarios.ts, src/simulator.ts, src/app.ts, and src/client/App.tsx.
- Node.js 20 or newer; Node.js 22 is recommended.
- npm, included with Node.js.
git clone https://github.com/m3yyyyy/cloud-ops-observability-lab.git
Set-Location .\cloud-ops-observability-lab
npm.cmd install
npm.cmd run devOpen http://127.0.0.1:5173. The Vite development server proxies API requests to the Express server on port 4100.
git clone https://github.com/m3yyyyy/cloud-ops-observability-lab.git
cd cloud-ops-observability-lab
npm install
npm run devOpen http://127.0.0.1:5173.
npm.cmd run build
npm.cmd startOpen http://127.0.0.1:4100.
Run the full local quality gate:
npm.cmd run checkOr run each stage separately:
npm.cmd run lint
npm.cmd test
npm.cmd run buildThe current suite contains 10 tests covering scenario generation, deterministic time series, incident correlation, error-budget exhaustion, health and snapshot endpoints, runbooks, invalid scenarios, and the absence of a write-capable scenario route.
The production build has also been checked in a real browser across all four scenarios and a 390-pixel mobile breakpoint, with no console errors or horizontal overflow.
All endpoints are read-only.
| Method | Route | Purpose |
|---|---|---|
GET |
/api/health |
Liveness and operating mode |
GET |
/api/scenarios |
Available deterministic scenarios |
GET |
/api/snapshot?scenario=healthy |
Complete typed dashboard snapshot |
GET |
/api/runtime |
Recent real runtime spans and logs |
GET |
/api/runbooks |
Local runbook index |
GET |
/api/runbooks/:id |
One local runbook |
GET |
/api/export?scenario=latency |
Download a structured JSON snapshot |
Valid scenario IDs are healthy, latency, error-spike, and dependency-outage. Unknown values return a structured 400 response.
Example incident fragment:
{
"title": "Checkout latency caused by inventory database contention",
"severity": "SEV-2",
"status": "investigating",
"likelyCause": "Row-lock contention in the inventory reservation query.",
"confidence": 0.91,
"traceIds": ["7f9c1b2a4d8e6c31"],
"alertIds": ["alert-inventory-latency", "alert-checkout-burn"]
}.
├── data/runbooks.json # local knowledge base
├── docs/ # screenshot and demo script
├── src/
│ ├── app.ts # read-only Express API
│ ├── server.ts # lifecycle and local binding
│ ├── telemetry.ts # OpenTelemetry SDK and span exporter
│ ├── simulator.ts # snapshot and correlation engine
│ ├── scenarios.ts # service topology and incident scenarios
│ ├── types.ts # shared data contracts
│ └── client/ # React dashboard and components
├── tests/ # simulator and API tests
├── Dockerfile # multi-stage production container
├── docker-compose.yml # hardened local container profile
├── render.yaml # optional Render deployment blueprint
└── .github/workflows/ci.yml # GitHub Actions quality gate
docker compose up --buildThen open http://127.0.0.1:4100. The Compose profile uses a read-only filesystem, a temporary /tmp, and no-new-privileges.
The included render.yaml builds the React assets, starts the Express service, binds to Render's network interface, and checks /api/health.
- Push the repository to GitHub.
- In Render, choose New → Blueprint.
- Select this repository and apply the detected blueprint.
- Open the service URL after the health check becomes green.
The hosted version remains a simulation; it does not gain access to your infrastructure.
- The server binds to
127.0.0.1by default. - No authentication secrets or environment credentials are required.
- No endpoint changes scenario state or executes a remediation action.
- No shell commands, arbitrary hosts, cloud APIs, or external collectors are exposed.
- Runtime telemetry is process-local and is cleared on restart.
- Error responses avoid leaking server internals.
See SECURITY.md before adapting the project to real telemetry.
- Add an optional OpenTelemetry Collector and Grafana profile behind an explicit Docker Compose flag.
- Add exemplars that link metric points directly to trace IDs.
- Add a human-approved incident annotation workflow with an append-only audit trail.
- Add configurable SLO policies and multi-window burn-rate alerts.
- Add deployment markers to the service timeline.
- Add persistence through a local SQLite adapter while keeping the default stateless.
- Add Playwright accessibility and visual-regression checks in CI.
Built a TypeScript and React observability control plane instrumented with OpenTelemetry, correlating service metrics, alerts, distributed traces, structured logs, SLO error budgets, and runbooks into explainable SEV-1/SEV-2 incident timelines. Shipped deterministic simulations, structured read-only APIs, 10 automated tests, GitHub Actions CI, responsive UI, Docker packaging, and safe local-first defaults.
- Why one typed snapshot keeps every signal consistent.
- How trace and span IDs connect logs to a distributed request path.
- Why error-budget burn is more actionable than availability alone.
- Where human approval would sit before any future remediation.
- Which safeguards would be required before accepting real production telemetry.
For a guided walkthrough, use docs/DEMO.md.
MIT © 2026 m3yyyyy

