Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

System Design Labs

Working through system design by building the concepts rather than reading about them.

Each lab takes a single idea, builds the smallest real thing that makes it concrete, breaks it on purpose, and records what actually happened — with numbers.

The subject in every lab is a real application: ea-qms-backend, a change control API for a pharmaceutical quality management system. Go, no frameworks, sqlc over hand-written SQL, PostgreSQL. Published as 20dumpling/ea-qms-backend.


Why build rather than read

The load balancer lab started as "put nginx in front of three servers" — three lines on a whiteboard. Getting there required understanding container networking, DNS resolution inside a bridge network, why localhost means something different inside a container, how a connection pool interacts with a database's global limit, and what a TCP handshake does when nothing answers.

None of that comes from the diagram. It comes from port collisions and exit codes.

The method:

build the smallest thing that makes the concept real
  → break it on purpose
  → measure what happened
  → write down what the failure taught

Breaking it is the load-bearing step. Anyone can make it work once.


Labs

Only sections that earn a build appear here. The curriculum also covers conceptual topics and paper design exercises — CAP, CDNs, "design Instagram" — which are studied but do not produce a lab folder. Forcing a build onto them would be busywork.

# Lab Concepts Status
01 Load Balancing Horizontal scaling · statelessness · L4 vs L7 · round-robin · reverse proxy · passive health checks · private networking ✅ Done
02 URL Shortener Base62 encoding · collision handling · read-heavy workload · where it hurts under load Next
03 ACID & Isolation Levels Dirty read · lost update · phantom · MVCC and why SERIALIZABLE fails rather than blocks Planned
04 Indexes & N+1 B-tree intuition · measured query time before and after · the write cost · N+1 with sqlc and its fix Planned
05 Database Replication Primary/replica · read-write split · replication lag · watching read-your-writes break Planned
06 Caching Cache-aside · TTL and invalidation · forcing a stampede on a hot key Planned
07 Queues & Idempotency Producer/consumer · at-least-once delivery · the double-effect · idempotency keys Planned
08 Rate Limiting Token bucket in middleware · per-client limits · API versioning Planned
09 Real-time WebSockets behind a load balancer · why sticky sessions appear · pub/sub fan-out Planned
10 Service Split & Saga One bounded context extracted · the operational cost · compensating transactions Planned
11 Resilience Patterns Timeouts · retry with backoff and jitter · circuit breaker · graceful degradation Planned
12 Security OWASP defenses · secrets handling · token validation boundaries Planned
13 Observability Structured logs · correlation IDs · metrics · tracing across services Planned
14 Payment Idempotency Exactly-once effects · audit trail · consistency under retry Planned
15 Cloud & Deployment Container orchestration · CI/CD · blue-green and canary releases Planned

Progress: 1 of 15.


Findings so far

01 — Load Balancing

  • Round-robin distributed 30 requests exactly 10/10/10. Strict rotation, not random — it is a counter, with no knowledge of which backend is busy.
  • The proxy hop cost ~1 ms measured like-for-like. Proxy overhead is almost never where latency lives.
  • Killing a replica mid-traffic produced zero client-visible failures but one request took 3.07 s against a ~4 ms baseline. That is the price of passive health checking: at least one request must fail before a dead backend is detected.
  • A token issued by one replica was accepted by the other two. Statelessness proven rather than assumed — though refresh tokens are stateful by design, which is what makes revocation possible.
  • Exhausting the database surfaced sorry, too many clients already at two of three replicas simultaneously. Connection pools are per-process; max_connections is global. The failure appears at the shared resource, not at the service that caused it.

Running any lab

Each lab folder has its own README with the full command sequence, the results, and a teardown. NOTES.md alongside it holds the line-by-line mechanics — TCP behaviour, directive semantics, the reasoning behind each default — for revision rather than for a first read.

Prerequisites are Docker and a local PostgreSQL. Labs are built with plain docker run rather than Compose, deliberately: every flag typed is a concept named, and Compose hides exactly the parts worth understanding.

About

Learning system design by building it — each lab takes one concept, builds the smallest real thing, breaks it on purpose, and records what happened with numbers.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors