Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion docs/background/waf/cost-optimization.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,9 @@ Cost Optimization is the ability to run workloads at the lowest price point that

## Applying this on {{brand}}

- Take advantage of per-second billing on virtual machines — only pay for the compute time actually used, and shut down non-production instances outside working hours.
- Take advantage of per-second billing on virtual machines.
Only pay for the compute time that actual work uses.
Shut down non-production instances outside working hours.
- Apply lifecycle and versioning policies on object storage to automatically move aging or infrequently accessed data into archival tiers.
- Use per-project quotas to prevent unplanned overspend and to attribute cost cleanly across teams, environments, or business units.
- Periodically re-evaluate the build-vs-offload trade-off:
Expand Down
9 changes: 5 additions & 4 deletions docs/background/waf/general-design-principles.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,10 +13,11 @@ These principles apply across all [five pillars](introduction.md) and should inf
- Match the deployment model to the workload's risk profile.
Use the {{brand_compliant}} for regulated or mission-critical data, the {{brand_public}} for general-purpose workloads, and the {{company}} Private Cloud where full isolation is required.
- Keep architectures open and portable.
{{brand}}'s OpenStack foundation avoids proprietary lock-in — design workloads so they remain portable across regions and, if ever needed, providers.
{{brand}}'s OpenStack foundation avoids proprietary lock-in.
Design workloads so they remain portable across regions and, if ever needed, providers.
- Build in security and compliance from day one, not as an afterthought before an audit.
- Test what you assume will work.
Regularly test backups, failover, and scaling — not just at launch.
Rather than testing only at launch, test backups, failover, and scaling regularly.
- Right-size continuously.
Cloud resources are elastic;
provisioning should be revisited as usage patterns change, not fixed at launch.
Cloud resources are elastic.
Provisioning should be revisited as usage patterns change, not fixed at launch.
30 changes: 19 additions & 11 deletions docs/background/waf/introduction.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
description: Cloud architecture is never finished — it is continuously evaluated, tested, and improved as your business and workloads evolve.
description: Cloud architecture is never finished. You continuously evaluate, test, and improve it as your business and workloads evolve.
---

# Introduction
Expand All @@ -18,8 +18,8 @@ It is organized around **five pillars:**

Each pillar defines design principles, practical guidance specific to {{brand}}'s OpenStack-based platform, and a self-assessment checklist you can use during design reviews.

The framework is deployment-model agnostic:
the same principles apply whether you run on {{brand_public}}, {{brand_compliant}}, or {{company}} Private Cloud, though the specific controls available differ between them (see the deployment model note below).
The framework is deployment-model agnostic: the same principles apply whether you run on {{brand_public}}, {{brand_compliant}}, or {{company}} Private Cloud.
The specific controls available differ between them (see the deployment model note below).

## Who this is for

Expand All @@ -30,23 +30,30 @@ the same principles apply whether you run on {{brand_public}}, {{brand_compliant
## How to use this framework

1. Define the workload.
Scope the review to one workload or system at a time — not your entire estate.
Scope the review to one workload or system at a time.
Do not scope your entire estate.
2. Work through each pillar.
Use the design principles as guidance and the checklist as a gap-finder.
3. Prioritize findings.
Not every gap needs fixing immediately — weigh risk, cost, and effort.
Not every gap needs immediate remediation.
Weigh the risk, cost, and effort.
4. Create an improvement backlog.
Track remediation items alongside your normal engineering backlog.
5. Repeat regularly.
Re-run the review after major changes, and at minimum quarterly for production workloads.

## A note on deployment models

{{company}} offers three IaaS deployment models built on the same open-source, OpenStack foundation:
Public Cloud (flexible, standard-tools environment for developers and SMBs), Compliant Cloud (enhanced security configuration, availability zones, and controls for regulated and mission-critical workloads), and Private Cloud (a dedicated, turnkey OpenStack environment).
{{company}} offers three IaaS deployment models built on the same open source, OpenStack foundation:

Many of the practices below — particularly under Security & Digital Sovereignty and Reliability — are strongest on Compliant Cloud or Private Cloud.
Choosing the right model is a foundational decision that should be made early, based on the workload's regulatory, security, and availability requirements, since it shapes which controls and practices in this framework are available to you from the outset.
- Public Cloud (flexible, standard-tools environment for developers and SMBs),
- Compliant Cloud (enhanced security configuration, availability zones, and controls for regulated and mission-critical workloads),
- Private Cloud (a dedicated, turnkey OpenStack environment).

Many practices are strongest on Compliant Cloud or Private Cloud, especially under Security & Digital Sovereignty and Reliability.
Choosing the right model is a foundational decision that you must make early.
Base it on the workload's regulatory, security, and availability requirements.
Your choice shapes which controls and practices in this framework are available from the outset.

## A note on {{brand}} Launch Pad

Expand All @@ -58,7 +65,8 @@ Run via OpenStack Heat, Ansible, or OpenTofu, it creates:
- a virtual router connected to an internal IPv4/IPv6 network with public internet access, and
- a Pad Ramp jump host that you can restrict to a specific source IP or network.

That's a genuinely useful starting point for reaching a brand-new environment.
Launch Pad does not create projects, quotas, IAM boundaries, a security baseline beyond the jump host's own access rule, logging, or multi-environment structure.
This is a genuinely useful starting point for reaching a brand-new environment.
Launch Pad does not create projects, quotas, or IAM boundaries.
Launch Pad does not create a security baseline beyond the jump host's own access rule, logging, or multi-environment structure.
Treat every other practice in this framework, including basic hygiene like project segmentation and security groups, as work that still needs to happen after Launch Pad hands off.

13 changes: 8 additions & 5 deletions docs/background/waf/operational-excellence.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,10 +18,12 @@ On {{brand}} this means treating infrastructure as code, standardizing deploymen

## Applying this on {{brand}}

- Use {{company}}'s Launch Pad (via Heat, Ansible, or OpenTofu) to bootstrap initial connectivity into a new environment — it creates an SSH keypair, a router and network with public access, and a restrictable jump host.
It's a convenient entry point, not a landing zone:
project structure, quotas, IAM boundaries, and a security baseline still need to be set up separately using the practices in this framework.
- Provision and manage resources through the {{gui}} (CCMP), CLI, or API — and drive routine provisioning through Infrastructure-as-Code rather than manual clicks, so changes are auditable and repeatable.
- Use {{company}}'s Launch Pad (via Heat, Ansible, or OpenTofu) to bootstrap initial connectivity into a new environment.
- It creates an SSH keypair, a router and network with public access, and a restrictable jump host.
- It is a convenient entry point, not a landing zone.
- Project structure, quotas, IAM boundaries, and a security baseline still need to be set up separately using the practices in this framework.
- Provision and manage resources through the {{gui}} (CCMP), CLI, or API.
- Drive routine provisioning through Infrastructure-as-Code rather than manual clicks, so changes are auditable and repeatable.
- Use OpenStack projects to cleanly separate development, test, and production environments, and to scope team access to only what each group needs.
- Using IaC as the standard for your infrastructure deployments.
It provides a consistent, standard methodology for development and deployment for all components of your workload.
Expand All @@ -33,7 +35,8 @@ On {{brand}} this means treating infrastructure as code, standardizing deploymen
## Self-assessment checklist

- Is infrastructure provisioned through code or API calls rather than ad hoc manual changes?
- If the project used {{company}}'s Launch Pad to bootstrap connectivity, has project structure, IAM, quotas, and a security baseline been set up separately — since Launch Pad itself does not provide these?
- If the project used {{company}}'s Launch Pad to bootstrap connectivity, have the project structure, IAM, quotas, and a security baseline been set up separately?
Launch Pad itself does not provide these.
- Do you have documented, tested runbooks for common operational events (failover, scale-out, incident response)?
- Are environments (dev/test/prod) cleanly separated using projects?
- Do deployments happen through a CI/CD pipeline rather than manual steps?
Expand Down
9 changes: 6 additions & 3 deletions docs/background/waf/performance-efficiency.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,17 +9,20 @@ Performance Efficiency is the ability to use computing resources efficiently to
## Design principles

- Select the compute, storage, and network resources that actually match the workload's characteristics.
- Use elasticity — scale resources up and down as demand changes rather than provisioning for peak permanently.
- Use elasticity.
Scale resources up and down as demand changes instead of provisioning for peak permanently.
- Prefer managed services over self-managed infrastructure where they remove undifferentiated operational work.
- Continuously monitor and benchmark rather than sizing once and forgetting.

## Applying this on {{brand}}

- Choose the VM flavor profile that fits the workload:
Generic for general-purpose use, Low Latency Disk for I/O-sensitive applications, or High Intensity CPU for compute-bound workloads.
- Right-size Kubernetes worker pools using mixed VM flavors through {{k8s_management_service}}-managed clusters, matching node types to the actual mix of workloads running on them.
- Right-size Kubernetes worker pools using mixed VM flavors through {{k8s_management_service}}-managed clusters.
Match node types to the actual mix of workloads.
- Match storage to access pattern: high-performance block storage for latency-sensitive workloads, S3-compatible object storage for scalable unstructured data, and archival/lifecycle tiers for cold data.
- Place workloads in the {{company}} region closest to your users — Stockholm, Karlskrona, or Frankfurt — to minimize latency.
- Place workloads in the {{company}} region closest to your users.
This minimizes latency.
- Monitor resource utilization through CCMP or the API on an ongoing basis, and adjust instance flavors, worker pool sizes, and volume types as usage patterns change rather than only at initial deployment.

## Self-assessment checklist
Expand Down
44 changes: 29 additions & 15 deletions docs/background/waf/reliability.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,8 @@ description: Cleura Compliant Cloud's architecture provides strong building bloc

Reliability is the ability of a workload to perform its intended function correctly and consistently, including its ability to recover quickly from failure.

{{brand_compliant}}'s multi-availability-zone architecture and built-in snapshot-based recovery give you strong building blocks here.
{{brand_compliant}} provides a multi-availability-zone architecture and built-in snapshot-based recovery.
These features give you strong building blocks.

## Design principles

Expand All @@ -18,41 +19,54 @@ Reliability is the ability of a workload to perform its intended function correc

## Applying this on {{brand}}

- Deploy production workloads across {{brand_compliant}}'s availability zones — each an independently powered and networked data center, interconnected by redundant, low-latency fiber — so a single facility failure does not take down the workload.
- Deploy production workloads across {{brand_compliant}}'s availability zones.
Each zone is an independently powered and networked data center, interconnected by redundant, low-latency fiber.
A single facility failure cannot take down the workload.
- Use {{company}}'s Load Balancer service to distribute traffic across instances and integrate it with auto-scaling for resilience under variable load.
- Enable {{company}}'s automated daily snapshots of virtual servers and storage volumes (retained for 10 or 30 days) as a built-in, low-effort disaster-recovery layer with fast RPO/RTO, without needing third-party backup tooling — and periodically test restoring from them.
There is also an option to choose immutable backups.
- Enable {{company}}'s automated daily snapshots of virtual servers and storage volumes.
The snapshots are retained for 10 or 30 days.
This built-in, low-effort disaster-recovery layer delivers fast RPO/RTO without third-party backup tooling.
Periodically test restoring from them.
You can also choose immutable backups.
- For workloads that need regional resilience beyond a single cloud region, architect cross-region disaster recovery across {{company}}'s regions (Stockholm, Karlskrona, Frankfurt).
- Use Managed Database and Managed Kubernetes/OpenShift services where you want {{company}} to handle patching, monitoring, and high-availability configuration on your behalf.

## Designing for reliability: Virtual Machines

VM-based workloads on {{brand}} should be architected so that no single instance, host, or availability zone is a point of failure.

- Use OpenStack anti-affinity server groups so replica instances of the same tier (e.g., web or application nodes) are automatically scheduled onto different physical hosts — a single hypervisor failure should never be able to remove every replica at once.
- Use OpenStack anti-affinity server groups to place replica instances of the same tier (e.g., web or application nodes) on different physical hosts.
A single hypervisor failure cannot remove every replica at once.
- Spread each tier's instances across multiple availability zones in the {{brand_compliant}}, not just across hosts within one zone, so a data-center-level event only affects part of your capacity.
- Boot instances from volumes rather than local ephemeral disk when the instance needs to survive host maintenance, evacuation, or rebuild without losing its state.
- Place a Load Balancer in front of any tier that has more than one instance, with health checks configured so unhealthy instances are automatically taken out of rotation rather than continuing to receive traffic.
- Enable the automated daily snapshot policy on both instances and attached volumes, and periodically run a real restore — not just a backup — to confirm recovery actually works within your target RTO.
- For stateful services such as databases or message queues, don't rely on snapshots alone if you need a low RPO — pair them with application-level replication (for example PostgreSQL streaming replication, MySQL/MariaDB replication, or MongoDB replica sets), or use Managed Database's own HA configuration where the workload fits a supported engine.
- Enable the automated daily snapshot policy on both instances and attached volumes.
Periodically run a real restore, not just a backup, to confirm recovery actually works within your target RTO.
- For stateful services such as databases or message queues, do not rely on snapshots alone if you need a low RPO.
- Pair them with application-level replication (PostgreSQL streaming replication, MySQL/MariaDB replication, or MongoDB replica sets).
Alternatively, use Managed Database's own HA configuration where the workload fits a supported engine.
- Be aware that {{brand}} does not provide a native, continuous VM-to-VM replication capability beyond the automated daily snapshot policy.
For general-purpose VMs where daily snapshots aren't a tight enough RPO, and where the application itself has no built-in replication, you will need to add replication yourself via general purpose backup/DR vendors.
- Treat instances as replaceable, not precious:
keep Infrastructure-as-Code up to date so a failed VM can be recreated automatically in minutes rather than repaired by hand.
- Monitor host and instance health continuously — through {{brand}} Managed Services or your own tooling — so degradation is caught before it becomes a customer-facing incident.
- Treat instances as replaceable, not precious.
- Keep Infrastructure-as-Code up to date so a failed VM can be recreated automatically in minutes instead of repaired by hand.
- Monitor host and instance health continuously through {{brand}} Managed Services or your own tooling.
Detect degradation before it becomes a customer-facing incident.

## Designing for reliability: containers and Kubernetes

For containerized workloads running on {{company}}'s managed Kubernetes ({{k8s_management_service}}), reliability is split between the platform layer, which {{company}} helps manage, and the workload layer, which is your responsibility to configure correctly.

- Spread worker node pools across multiple availability zones — {{k8s_management_service}} support multi-AZ pools — so that losing one zone only removes part of your worker capacity rather than the whole cluster.
- Spread worker node pools across multiple availability zones.
{{k8s_management_service}} supports multi-AZ pools.
Losing one zone removes only part of your worker capacity instead of the whole cluster.
- Let {{k8s_management_service}} manage control-plane high availability (etcd and the API server) rather than self-managing it, and put your reliability effort into the workload layer where you have the most control.
- Set Pod anti-affinity rules or topology spread constraints so that replicas of the same service are scheduled onto different nodes and, ideally, different availability zones — the default scheduler will happily stack all replicas on one node otherwise.
- Set Pod anti-affinity rules or topology spread constraints so that replicas of the same service land on different nodes and, ideally, different availability zones.
The default scheduler stacks all replicas on one node otherwise.
- Define readiness and liveness probes on every Deployment so Kubernetes automatically stops routing traffic to pods that aren't ready and restarts pods that have hung.
- Use PodDisruptionBudgets so voluntary disruptions — node upgrades, cluster autoscaler scale-down, routine maintenance — can't remove every replica of a service at the same time.
- Use PodDisruptionBudgets so voluntary disruptions like node upgrades, cluster autoscaler scale-down, and routine maintenance cannot remove every replica of a service at the same time.
- Enable the cluster autoscaler on worker pools so node capacity grows automatically under load and shrinks safely afterward, paired with the Horizontal Pod Autoscaler at the workload level for pod-level scaling.
- Remember that block storage-backed PersistentVolumeClaims are zone-affine:
a StatefulSet's volume ties its pod to the availability zone where that volume was created.
- Remember that block storage-backed PersistentVolumeClaims are zone-affine.
A StatefulSet's volume ties its pod to the availability zone where that volume was created.
Plan storage-backed workload placement with this constraint in mind, particularly when combining it with multi-AZ node pools.
- Run your ingress controller with at least two replicas behind the Load Balancer service, with its own health checks, so the ingress layer itself isn't a single point of failure.
- Use rolling update strategies (maxUnavailable/maxSurge) for deployments and node pool upgrades, and rehearse upgrades in a non-production cluster before rolling them out to production.
Expand Down
Loading
Loading