Skip to content

Repository files navigation

Cloud Ops Observability Lab

CI Live demo Node.js OpenTelemetry Safety License

A portfolio-grade, local-first observability control plane that turns service metrics, alerts, distributed traces, structured logs, and SLO error budgets into one explainable incident story.

Live demo: cloud-ops-observability-lab.onrender.com — the free instance may need up to 50 seconds to wake after inactivity.

The project demonstrates modern platform engineering without requiring a cloud account or touching production infrastructure. Four deterministic scenarios let reviewers explore healthy operations, latency, authentication failures, and a dependency outage safely.

Latency regression dashboard showing service health, alerts, and incident correlation

Correlated alerts, incident evidence, timeline, and distributed trace waterfall

Why this project

Modern IT operations is moving beyond isolated dashboards toward correlated, machine-readable telemetry that humans and agents can reason about. This lab demonstrates that workflow while keeping the important operational boundary explicit: observation and explanation are allowed; infrastructure changes are not.

It is designed to show employers that the author can connect software engineering and IT operations:

  • instrument a real Node.js service with the OpenTelemetry SDK;
  • model metrics, traces, logs, alerts, incidents, timelines, and SLOs with TypeScript;
  • correlate evidence across services instead of presenting unrelated charts;
  • design structured JSON APIs for automation and AI-assisted operations;
  • build an accessible responsive React dashboard;
  • ship tests, CI, containers, deployment metadata, security notes, and documentation.

Highlights

  • Five-service topology with owners, versions, regions, dependencies, health states, and golden signals.
  • Four repeatable scenarios: healthy baseline, latency regression, authentication errors, and dependency outage.
  • Cross-signal incident correlation with severity, likely cause, confidence, evidence, affected services, trace links, alert links, and a timeline.
  • Distributed trace waterfall with parent-child spans, attributes, duration, and error status.
  • Structured log explorer with level filters, text search, trace IDs, span IDs, and attributes.
  • SLO and error-budget view with a 30-day window, remaining budget, burn rate, and policy state.
  • Real runtime telemetry captured by the official OpenTelemetry Node SDK for Express requests.
  • Local runbook library exposed through read-only endpoints.
  • JSON export for downstream automation or incident-analysis tools.
  • Safe by default: no credentials, cloud calls, shell execution, write actions, or remediation endpoints.

Scenario matrix

Scenario Customer impact Signals to inspect Correlated outcome
Healthy baseline None Stable golden signals and budgets No incident or alert
Latency regression Slow checkout Inventory latency, checkout burn, warning logs, slow database span SEV-2 with 91% confidence
Authentication errors Some sign-ins fail Identity error rate, edge 401s, signing-key errors SEV-2 with 88% confidence
Dependency outage Checkout and notifications disrupted Availability collapse, retries, failed spans, exhausted budgets SEV-1 with 96% confidence

Architecture

flowchart LR
    U[React dashboard] -->|GET scenario snapshot| A[Express API]
    A --> V[Zod validation]
    V --> S[Deterministic scenario engine]
    S --> M[Metrics and SLO model]
    S --> C[Incident correlation]
    S --> T[Simulated distributed traces]
    S --> L[Structured scenario logs]
    A --> R[Local runbook JSON]

    A --> O[OpenTelemetry Node SDK]
    O --> E[In-memory span exporter]
    E --> X[Runtime telemetry view]

    M --> J[Typed observability snapshot]
    C --> J
    T --> J
    L --> J
    R --> J
    X --> J
    J --> U

    style U fill:#10233a,stroke:#58a6ff,color:#fff
    style J fill:#12352d,stroke:#42d392,color:#fff
    style O fill:#2d2552,stroke:#9d8cff,color:#fff
Loading

The dashboard performs one snapshot request per refresh. Every panel renders from that same typed response, preventing the metrics, incident narrative, trace, logs, and SLO state from contradicting one another.

What is real and what is simulated

Area Implementation
Express request spans Real OpenTelemetry SDK spans with non-zero trace and span IDs
Request logs Real structured JSON runtime logs in a bounded memory buffer
Service metrics and SLO history Deterministic local scenario data
Distributed service traces Deterministic cross-service examples for investigation practice
Runbooks Bundled local JSON; never fetched or executed remotely
Infrastructure actions Intentionally absent

Beginner-friendly stack

Technology Why it was chosen
TypeScript One language and shared types across the server, simulator, tests, and UI
Express 5 Small, readable API surface with familiar HTTP concepts
React 19 + Vite Fast local feedback and component-based dashboard development
OpenTelemetry JS Vendor-neutral observability concepts and official runtime instrumentation
Zod Clear validation at the API boundary
Node test runner Reliable tests without a large additional test framework

You do not need prior experience with these tools to run the project. The first useful files to read are src/scenarios.ts, src/simulator.ts, src/app.ts, and src/client/App.tsx.

Quick start

Requirements

  • Node.js 20 or newer; Node.js 22 is recommended.
  • npm, included with Node.js.

Windows PowerShell

git clone https://github.com/m3yyyyy/cloud-ops-observability-lab.git
Set-Location .\cloud-ops-observability-lab
npm.cmd install
npm.cmd run dev

Open http://127.0.0.1:5173. The Vite development server proxies API requests to the Express server on port 4100.

macOS or Linux

git clone https://github.com/m3yyyyy/cloud-ops-observability-lab.git
cd cloud-ops-observability-lab
npm install
npm run dev

Open http://127.0.0.1:5173.

Production-style local run

npm.cmd run build
npm.cmd start

Open http://127.0.0.1:4100.

Verification

Run the full local quality gate:

npm.cmd run check

Or run each stage separately:

npm.cmd run lint
npm.cmd test
npm.cmd run build

The current suite contains 10 tests covering scenario generation, deterministic time series, incident correlation, error-budget exhaustion, health and snapshot endpoints, runbooks, invalid scenarios, and the absence of a write-capable scenario route.

The production build has also been checked in a real browser across all four scenarios and a 390-pixel mobile breakpoint, with no console errors or horizontal overflow.

API

All endpoints are read-only.

Method Route Purpose
GET /api/health Liveness and operating mode
GET /api/scenarios Available deterministic scenarios
GET /api/snapshot?scenario=healthy Complete typed dashboard snapshot
GET /api/runtime Recent real runtime spans and logs
GET /api/runbooks Local runbook index
GET /api/runbooks/:id One local runbook
GET /api/export?scenario=latency Download a structured JSON snapshot

Valid scenario IDs are healthy, latency, error-spike, and dependency-outage. Unknown values return a structured 400 response.

Example incident fragment:

{
  "title": "Checkout latency caused by inventory database contention",
  "severity": "SEV-2",
  "status": "investigating",
  "likelyCause": "Row-lock contention in the inventory reservation query.",
  "confidence": 0.91,
  "traceIds": ["7f9c1b2a4d8e6c31"],
  "alertIds": ["alert-inventory-latency", "alert-checkout-burn"]
}

Project structure

.
├── data/runbooks.json             # local knowledge base
├── docs/                          # screenshot and demo script
├── src/
│   ├── app.ts                     # read-only Express API
│   ├── server.ts                  # lifecycle and local binding
│   ├── telemetry.ts               # OpenTelemetry SDK and span exporter
│   ├── simulator.ts               # snapshot and correlation engine
│   ├── scenarios.ts               # service topology and incident scenarios
│   ├── types.ts                   # shared data contracts
│   └── client/                    # React dashboard and components
├── tests/                         # simulator and API tests
├── Dockerfile                     # multi-stage production container
├── docker-compose.yml             # hardened local container profile
├── render.yaml                    # optional Render deployment blueprint
└── .github/workflows/ci.yml       # GitHub Actions quality gate

Docker

docker compose up --build

Then open http://127.0.0.1:4100. The Compose profile uses a read-only filesystem, a temporary /tmp, and no-new-privileges.

Optional Render deployment

The included render.yaml builds the React assets, starts the Express service, binds to Render's network interface, and checks /api/health.

  1. Push the repository to GitHub.
  2. In Render, choose New → Blueprint.
  3. Select this repository and apply the detected blueprint.
  4. Open the service URL after the health check becomes green.

The hosted version remains a simulation; it does not gain access to your infrastructure.

Security and operating boundaries

  • The server binds to 127.0.0.1 by default.
  • No authentication secrets or environment credentials are required.
  • No endpoint changes scenario state or executes a remediation action.
  • No shell commands, arbitrary hosts, cloud APIs, or external collectors are exposed.
  • Runtime telemetry is process-local and is cleared on restart.
  • Error responses avoid leaking server internals.

See SECURITY.md before adapting the project to real telemetry.

Roadmap

  • Add an optional OpenTelemetry Collector and Grafana profile behind an explicit Docker Compose flag.
  • Add exemplars that link metric points directly to trace IDs.
  • Add a human-approved incident annotation workflow with an append-only audit trail.
  • Add configurable SLO policies and multi-window burn-rate alerts.
  • Add deployment markers to the service timeline.
  • Add persistence through a local SQLite adapter while keeping the default stateless.
  • Add Playwright accessibility and visual-regression checks in CI.

Resume-ready description

Built a TypeScript and React observability control plane instrumented with OpenTelemetry, correlating service metrics, alerts, distributed traces, structured logs, SLO error budgets, and runbooks into explainable SEV-1/SEV-2 incident timelines. Shipped deterministic simulations, structured read-only APIs, 10 automated tests, GitHub Actions CI, responsive UI, Docker packaging, and safe local-first defaults.

Demo talking points

  • Why one typed snapshot keeps every signal consistent.
  • How trace and span IDs connect logs to a distributed request path.
  • Why error-budget burn is more actionable than availability alone.
  • Where human approval would sit before any future remediation.
  • Which safeguards would be required before accepting real production telemetry.

For a guided walkthrough, use docs/DEMO.md.

License

MIT © 2026 m3yyyyy

About

Local-first OpenTelemetry observability lab that correlates metrics, traces, logs, alerts, SLOs and incident timelines.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages