My personal Kubernetes platform, in the open. Everything the cluster runs is described as files in this repository, and changes go live by being merged here rather than by anyone running commands against the cluster — that pattern is called GitOps, and Flux is what applies it.
This is a working system rather than a product: it is shaped around what I run, and it is not
packaged for reuse. Look around anyway — if you are building something similar, the repository
layout and the guides in docs/ are the useful parts, and
docs/TEMPLATING.md lists exactly what a fork has to change. 🙌🏻
For local development:
- Docker — runs the cluster on your machine.
- KSail — creates and manages that cluster. CI uses the same tool to validate and security-scan manifests, but does not boot a cluster; running one is a local step.
For the production cluster:
- Hetzner Cloud — Infrastructure provider and managed Cloud Load Balancer for cluster ingress. KSail's native Hetzner provider handles Talos boot, CCM, CSI, and kubeconfig.
- Cloudflare — DNS (A/AAAA records pointed at the Hetzner Cloud Load Balancer) and Origin CA.
- Flux GitOps — applies the manifests to the cluster. It does not read this repository directly: the manifests are packaged and published as an OCI artifact, and Flux pulls that.
- SOPS and Age — encrypt the starting
secrets that are committed here (the
*.enc.yamlfiles), so they can live in Git safely. - OpenBao and the External Secrets Operator —
where most secrets live once the cluster is running. Selected values from the encrypted files
above are seeded into OpenBao, and the operator copies each one into the namespace that needs it.
The rest — the bootstrap-critical values, plus Dex and oauth2-proxy — are still substituted
straight into manifests by Flux at apply time, so OpenBao is not the runtime path for everything
committed here. Full picture:
docs/secret-rotation.md.
Important
The committed secrets are encrypted with my keys, so you cannot decrypt them. To run this locally or in your own Hetzner project, swap in your own:
- Fork this repo.
- Create your own Age keys.
- Update
.sops.yamlin the repository root. - Put your Age key in your fork's GitHub secrets.
- Re-encrypt every
*.enc.yamlfile underk8s/with your own keys.
To run this cluster locally, simply run:
ksail cluster create
ksail workload push
ksail workload reconcilePorts 80 and 443 are mapped to localhost for you (via extraPortMappings in ksail.yaml),
and the hosts file points the *.platform.lan names at 127.0.0.1 — so deployed services
open in a browser.
The local cluster is a thin test-bed: somewhere to try one component before promoting it to production, not a copy of production. It starts with core infrastructure only — networking and gateway, DNS, TLS, Flux, policy, vertical pod autoscaling, secrets, single sign-on, and the CloudNativePG operator. Note that the operator is installed but no database is created: apps are opt-in, so a plain local bring-up has no PostgreSQL instance running.
Everything heavier (observability, request-rate autoscaling, backup, runtime security, the VM stack,
…) and all apps are opt-in. Uncomment what you want in these files — each carries a copy-paste
template — then re-run ksail workload push && ksail workload reconcile:
k8s/providers/docker/infrastructure/controllers/kustomization.yamlk8s/providers/docker/infrastructure/kustomization.yamlk8s/providers/docker/apps/kustomization.yaml
To tear down:
ksail cluster deleteBefore pushing, validate manifests with schema-aware checks and Flux variable substitution:
# Validate local cluster manifests (default)
ksail workload validate
# Validate prod cluster manifests
ksail --config ksail.prod.yaml workload validateThis is faster than a full cluster test and catches YAML errors, missing fields, and broken kustomize overlays.
CI also checks that the publish-workflow signing-revision report discovers exactly the consumers declared by the production overlay and its Flux layers:
bash scripts/guard-consumer-discovery-conservation.sh
bash scripts/tests/test-guard-consumer-discovery-conservation.shThis check uses yq, realpath and kubectl kustomize without a cluster or credentials.
The overlay and all production roots render from an isolated snapshot of KSail's published
k8s files: YAML/YML/JSON files selected by case-insensitive extension, without
ignore-file filtering. Directory links are not traversed and other inputs are omitted; selected file links
are read only after their targets are bounded to the source tree and staged as regular bytes.
Reachable resources, bases, components and file-loader inputs (patches, transformer
configurations, CRDs, generators, replacements and OpenAPI definitions) must resolve within
that artifact. Absolute, nonlocal, omitted or uninspectable inputs are unknown and fail before
rendering. A separate declared-path count must match the complete extracted path census;
successful but empty or truncated reader output cannot authorize a build. Inline patches,
single-document transformer configs and generator literals remain supported. Inline generator
and transformer strings must have exactly one document according to the YAML parser;
decoded bytes must match the original serialized string, and multiple documents or incomplete
parser receipts are unknown. Separators inside literal data
remain supported.
Contained relative manifest bases and selected file links remain supported; unused Kustomizations
are not traversed. Empty regular selected files and selected directory links prevent publication
and fail the check. Empty selected link targets are conservatively refused as unknown.
It compares
literal OCIRepository names and namespaces, exact URLs, effective refs and signer subjects.
An overlay rename is a divergence even when it keeps the same artifact and revision.
Identical consumer rows repeated across roots are deduplicated. All rendered top-level Flux
OCI declarations are checked for conflicting contracts that could overwrite an attributed
consumer, including unsigned sources; an absent namespace cannot establish disjointness.
Such conflicts are unknown and fail the check. Unrelated unsigned identities remain outside
the report. Guard-only identity rows preserve the
report's four-column --list-consumers output and artifact-based revision tables. It binds the
layers to the platform artifact configured in ksail.prod.yaml and the generated
OCIRepository/flux-system/flux-system source. Agreeing references to a different source, or
declarations that redirect that source, are unknown and fail the check. FluxInstance source
patches are supported only when their operations preserve the source identity and artifact
(verification and ref fields). Explicit source content selectors (ignore and layerSelector)
are refused because the same URL can produce a different tree. A production root's non-empty
targetNamespace is also an unseen Flux transform and is refused.
Admission mutations that can match consumer or production-root kinds are refused, as are
unbounded kind matches and unevaluated CEL mutations. Literal mutations of unrelated kinds
remain supported. Mapping-backed controller templates cannot add a Kustomization on the
platform source: its additional path would be missing from the rendered set. A declared
semverFilter also fails closed because the signing-revision resolver does not evaluate
Flux's tag filtering. Only objects in the Flux source API group count as OCI consumers,
and an attributed source must have no suspension field or a literal boolean false.
Generated sources, roots and controller carriers are refused; literal unrelated generated
kinds remain supported. Consumer-producing kro instances are counted in mapping-backed
controller carriers as well as top-level documents. Substituted structural or consumer
contract keys are unknown, just like substituted contract values. Nested source
references must also be literal, including an inherited namespace. An OCI-producing
kro definition may retain dormant references only while its complete literal schema
GVK has zero matching instances across every production render. Native mutating
webhooks that can target sources, roots or their policy/controller carriers are
refused because the static build cannot evaluate their callbacks; literal unrelated
resource groups and Pod-only rules remain supported.
The same boundary includes policies and controller carriers that can create or
rewrite those resources later. Nested ResourceSet string templates and substituted
carrier keys are unknown. Direct and foreach clone lists are checked against every
consumer-producing kro schema across all production roots; only literal foreign
group/version selectors can rule out a schema. CEL generation remains unevaluated.
Mapping-backed ResourceSet resources and Kyverno generation branches must name
literal object types before that comparison. Controller template expressions in
nested source references are also unknown; workload metadata references remain
supported when the generated object type is literal.
Nested FluxInstance templates also create sources and roots outside the static
render and are refused. Dormant OCI-producing kro definitions retain the complete
schema and zero-instance requirement across every root.
Flux substitutes the final YAML after the build, so a variable in an object or template kind or API version can hide an OCI consumer from literal discovery. The check refuses those object types, top-level source name/namespace variables, and consumer URL, ref or signer-subject variables; ordinary variables in workload fields, ConfigMap data and registry credentials remain supported. A same-named source with an omitted or empty namespace is checked as a possible platform source; an explicit different tenant namespace keeps its own source contract. A failed reader or render is unknown even when it produced partial output. Tenant artifacts and Helm chart output remain outside this repository's static comparison.
Local development cluster running on Docker via KSail. Uses Talos with the Docker provider. A small, thin manual test-bed (see Usage) — bring up a component, try it, then promote it to prod.
- 1 control-plane node + 1 worker node (Docker containers)
- Config:
ksail.yaml
Cloud cluster running on Hetzner Cloud via KSail's native Hetzner provider. Deployed via v* tags through the CD pipeline, and validated in the merge queue via the CI pipeline.
- 3× Hetzner CX33 control planes + 3× CX33 static workers + autoscaling (x86 4 vCPU 8GB RAM 80GB SSD each)
- Config:
ksail.prod.yaml
A high-level inventory of what Flux reconciles onto the cluster. The manifests live under k8s/bases/infrastructure/ and k8s/bases/apps/, with provider-specific pieces (Hetzner CCM/CSI, Longhorn, external-dns, …) under k8s/providers/. The exact set is overlay-dependent: the Hetzner/prod overlay deploys the full base set described below, while the local (docker) overlay is a thin manual test-bed that ships only the core controllers by default and makes the rest opt-in.
Infrastructure
- GitOps & config — Flux Operator, Reloader
- Networking — Cilium (CNI + Gateway API), CoreDNS, external-dns (Cloudflare), Hetzner CCM (prod)
- Certificates — cert-manager, trust-manager, Cloudflare Origin CA issuer
- Secrets — OpenBao + External Secrets Operator (runtime), SOPS + Age (at-rest seeds)
- Identity / SSO — Dex (OIDC) with oauth2-proxy / auth-proxy; see
docs/oidc-kubectl.md - Policy & runtime security — Kyverno, Kubescape, Tetragon; see
docs/runtime-security.md - Storage — Longhorn (replicated block/RWX, prod via Hetzner CSI), CloudNativePG (PostgreSQL operator); see
docs/rwx-storage.md - Autoscaling — Cluster Autoscaler (nodes), SIG Descheduler (pod rebalancing + node consolidation), Vertical Pod Autoscaler, KEDA request-rate autoscaling (homepage/umami); see
docs/node-autoscaling.md - Progressive delivery — Flagger (Gateway API canary deployments with SLO-gated automated rollback, metrics from Coroot); see
docs/progressive-delivery.md - Observability — Coroot (self-hosted, eBPF: metrics, logs, traces, profiling, service map, SLO alerting, cost allocation); see
docs/dr/alerting.md - Backup / DR — Velero with CloudNativePG backups; see
docs/dr/ - Virtualization — KubeVirt + CDI (local/CI only; disabled on the Hetzner/prod overlay)
- Testing — Testkube (local/CI only; not deployed to prod)
Apps (k8s/bases/apps/)
- Homepage (dashboard), Headlamp (Kubernetes UI), Umami (analytics), Actual Budget (budgeting),
whoami(debug) - FleetDM (device management) is parked — disabled 2026-06-03; its manifests are retained for re-enabling (see
k8s/bases/apps/kustomization.yaml) - Tenants — apps deployed from their own repositories (
ascoachingogvaner,wedding-app); seedocs/TENANTS.md
The cluster uses Flux GitOps to reconcile the state of the cluster with the single source of truth stored in this repository and published as an OCI image. KSail is used for local development, CI/CD testing, and production deployments. For prod, nodes are provisioned on Hetzner Cloud by KSail's native Hetzner provider, which also installs the Hetzner CCM and CSI drivers.
All environments use the Talos Kubernetes distribution. Local development and CI use Talos with the Docker provider; prod uses Talos with the Hetzner provider.
The cluster configuration is stored in the k8s/* directories where the structure is as follows:
clusters/: Contains the cluster specific configuration for each environment.providers/: Contains the provider specific configuration.bases/: Contains the different bases that are used for the different clusters and providers.infrastructure: Contains the different infrastructure components that are used for the different clusters and providers.apps: Contains the different apps that are used for the different clusters and providers.bootstrap: The foundational bootstrap layer (renamed fromvariables/). Holds the shared substitution variables (variables-baseConfigMap + SOPS-encrypted Secret) and cluster-scoped PriorityClasses (e.g.platform-critical), reconciled by thebootstrapFlux Kustomization before everything thatdependsOnit.
Two things stack up. First, each environment points at a provider, which patches the shared resources — so a change can be made for one cluster, for one provider, or for everything at once:
graph LR
subgraph "Cluster-specific"
local["clusters/local"]
prod["clusters/prod"]
end
subgraph "Provider-specific"
docker["providers/docker"]
hetzner["providers/hetzner"]
end
subgraph "Shared"
bases["bases/*"]
end
local --> docker
prod --> hetzner
docker --> bases
hetzner --> bases
Second, Flux applies those layers in order, each waiting for the one before it to come up:
graph TB
bootstrap["bootstrap"]
controllers["infrastructure-controllers"]
infra["infrastructure"]
apps["apps"]
controllers -- "depends on" --> bootstrap
infra -- "depends on" --> controllers
apps -- "depends on" --> infra
Each layer in that chain has a matching folder under providers/<provider-name>/ and bases/. The
infrastructure layer, for example, is backed by k8s/providers/<provider-name>/infrastructure/ and
k8s/bases/infrastructure/.
The layer definitions themselves are written once in k8s/clusters/base/ with placeholders where the
cluster and provider names go; each cluster's overlay fills those in. Only the per-cluster
bootstrap/ directory holds manifests unique to one cluster.
docs/TEMPLATING.md— the exact files a fork must edit, including how those placeholders get filled in.docs/TENANTS.md— adding a tenant: an app that runs on the platform from its own repository.
Deeper guides and design notes live in docs/:
TEMPLATING.md— the exact set of files a fork needs to edit to stand up its own instance.TENANTS.md— onboarding a new GitOps tenant (an app that runs on the platform from its own repository).deletion-and-data-retention.md— the decision that Git owns deletion and the storage layer keeps the data, and the order it rolls out in.node-autoscaling.md— how the Cluster Autoscaler is configured on Hetzner.oidc-kubectl.md— authenticatingkubectlagainst the cluster via OIDC.progressive-delivery.md— Flagger Gateway API canary deployments and the per-app onboarding recipe.runtime-security.md— Tetragon runtime security.rwx-storage.md— Longhorn replicated / RWX storage.secret-rotation.md— the secrets architecture (SOPS → OpenBao → External Secrets) and rotation design.dr/— disaster-recovery runbooks (backup/restore drills, OpenBao crypto custody, Velero + CloudNativePG, alerting).
Note
Prices are approximate and may be outdated.
| Item | No. | Per unit | Total in Actual | Total in $ |
|---|---|---|---|---|
| Cloudflare Domains | 2 | $0,87 | $1,74 | $1,74 |
| Hetzner CX33 (prod) | 6 | €6,49 | €38,94 | $44,21 |
| Hetzner Cloud LB LB11 (prod) | 1 | €5,39 | €5,39 | $6,12 |
| Total | $52,07 |