HA HashiCorp Vault: Single Source of Secrets on RKE2 Homelab (The secured Safe)
Episode 6 of the homelab build series. The storage floor is down, the filing cabinets are full, and every one of them has a key. Those keys need somewhere secure to live, and this episode is that secured safe: a highly a
Episode 6 of the homelab build series. The storage floor is down, the filing cabinets are full, and every one of them has a key. Those keys need somewhere secure to live, and this episode is that secured safe: a highly available, three-node HashiCorp Vault cluster, built and configured with Terraform, delivering secrets to pods through the Vault Secrets Operator with VaultAuth and VaultStaticSecret, and VaultDynamicSecret ready for short-lived credentials.
The Vault UI at vault.georgehomelab.com, signed in through Keycloak, on the Raft status page: one active node, two standbys, Seal Type: transit.
SSO Login with keycloak and Oauth2-proxy using google and Github social login.
Introduction
The cluster can survive a node failure. Could my secrets?
Everything in this homelab is built to survive. Five control-plane nodes. Three datacenters. Ceph keeping three copies of every byte. Postgres clusters that fail over on their own. Lose a node, a disk, even a whole datacenter, and the platform keeps running.
But every one of those systems is holding a key. Tekton Chains holds the private key that signs every image I build. Kyverno checks every image against its public half. Harbor hands out robot credentials. Every database has a password. So underneath the whole series sat a question I kept stepping around: where do the keys to the kingdom actually live, and what happens when that place goes down?
For a while, the honest answer was gitignored files and kubectl create secret run from my laptop. A pile of credentials wearing a security costume.
So I built a safe. A highly available HashiCorp Vault: three nodes on Raft, one active and two standbys, so any one of them can die and nobody notices. Terraform writes every secret. The Vault Secrets Operator delivers them to pods. No passwords in git, ever. I was proud of it.
Then I found out it had been dead for two months.
One token, the one the main Vault uses to fetch its unseal key from a second, tiny Vault, had quietly run out. A comment in my own playbook promised it renewed itself. It didn't. The main Vault couldn't unlock itself, and every pod went round in circles. Nothing visibly broke, because the secrets already handed out kept working. That's the scary part: a dead safe can look exactly like a healthy one.
So this episode isn't about installing Vault and declaring victory. It's about making the safe as reliable as everything it protects. Now, when the whole lab loses power, it opens itself in about two minutes, all three nodes healthy, with no human involved. If it ever stays shut, an alert tells me.
Because a secure platform isn't one where secrets are merely hidden. It's one where every app gets exactly the secrets it needs, without ever having to own them.
Why a safe at all
There are two kinds of homelab: the ones that have leaked a secret into a git repo, and the ones that are about to. I've been both.
Every credential starts life somewhere careless. A password in a values.yaml. An API key in a .env that's "definitely gitignored". A kubectl create secret run from a laptop at 1 a.m. that nobody can trace any more. It all works right up until you need to rotate something, rebuild the cluster, or work out how many copies of a password exist.
What I wanted fits in one sentence: every credential lives in one place, every workload fetches it the same way, and rotating one is a terraform apply.
Here's what that replaced:
- Secrets in git, even encrypted ones, which always end up in someone's terminal history or a screenshot eventually.
- Passwords in Helm values, which end up in the rendered manifests and the release history.
-
kubectl create secretfrom a laptop, which is fine until anyone else has to rebuild what you did.
Vault, or something lighter?
"Just use SOPS" is a fair question, so here's my fair answer. These tools don't all solve the same problem:
- Encrypt and commit (Sealed Secrets, SOPS): the secret lives in git, encrypted. They answer "how do I put a secret in git safely?"
- Sync from somewhere else (External Secrets Operator, Vault Secrets Operator): delivery trucks, not warehouses. They copy secrets from a store into Kubernetes.
- A secrets engine (Vault, cloud secret managers): something that can store, issue, rotate, revoke and audit credentials, and do encryption as a service.
| Secret lives in | Short-lived creds | Rotation | Audit log | Encryption as a service | Runtime dependency | |
|---|---|---|---|---|---|---|
| Vault | Vault | β |
terraform apply, then re-sync |
β | β | Must be up and unsealed |
| Sealed Secrets | git, encrypted | β | re-encrypt and commit | β | β | its controller |
| SOPS + age | git, encrypted | β | re-encrypt and commit | β | β | none |
| External Secrets | whatever backend | backend's | backend's | backend's | β | backend plus operator |
| Cloud secret manager | the cloud | β | provider API | β | β | the cloud |
If you've got a handful of static secrets and one GitOps loop, use Sealed Secrets or SOPS and stop there. If you live in one cloud, use its secret manager. I picked Vault for two reasons the git-based tools can't touch. Its transit engine is what lets one Vault unseal another, and the same Vault answers to pods, to Terraform and to me, with every read logged.
The cost is the one the rest of this article keeps coming back to. Vault is a live, stateful dependency. If it's sealed or down, nothing new gets a secret. SOPS can't fail that way, and I learned exactly how much that matters.
Two Vaults: the safe and the key-holder
Vault is sealed at rest. When a Vault pod starts, it can read its data from disk but can't decrypt any of it until something supplies the root key. In a cloud you'd hand that job to the cloud's key service. I don't have one, so I built a small one.
The two-Vault architecture: a three-pod main Vault that cannot unseal itself, and the one-pod transit Vault that holds the key. Links to the interactive diagram.
Click the image for the interactive version. Direct link: INTERACTIVE-DIAGRAM-URL
-
The main Vault runs three pods in the
vaultnamespace (Vault 1.21.2, integrated Raft storage). Its config has aseal "transit"block. On start it sends its encrypted root key to the transit Vault, gets it back decrypted, and unseals itself. -
The transit Vault is one pod in
vault-transit. It holds a single transit key and nothing else, and it's sealed the old-fashioned way, with one Shamir share.
The transit Vault is the single point of failure for the whole secret tier, and I chose that on purpose. The alternative is Shamir on the main Vault, which means unsealing three pods by hand after every restart. One small Vault is easier to look after, as long as something actually looks after it. That turned out to be the hard part.
kubectl get pods -n vault next to kubectl get pods -n vault-transit: three main pods, the transit pod, and its unsealer.
The main Vault
Each of the three main pods has a 10Gi Raft data volume and a 5Gi audit volume, both on ceph-block. One pod is active and the other two are standbys that forward writes to it. If the active pod dies, Raft elects a new leader from the other two.
Today they're spread one per datacenter: dc1-app-d, dc2-app-c and dc3-app-d. Mind you, that's luck, not policy, and I'll come back to it at the end.
That's what HA buys you here. Kill the active pod and one of the standbys takes over the leadership, while the Service keeps sending requests to whichever pods are up. Reboot a node and its Vault pod comes back, unseals itself against transit, and rejoins the cluster from Raft. At no point does a human type a key.
The UI is at vault.georgehomelab.com behind Istio, and you log in through Keycloak. Yes, Vault is gated by Keycloak while Keycloak's database password lives in Vault. The first deploy breaks that circle with the root token, and only after Keycloak is up does the OIDC login get configured.
You can see the login page above. The Vault UI on the Raft page: one active node, two standbys, Seal Type: transit, Recovery Seal Type: shamir.
AUTOMATING HASHICORP VAULT CONFIGURATION WITH TERRAFORM
Vault Config β prod Environment
Calls the vault-config module against the homelab Vault. One module call, one terraform apply per environment.
terraform/vault-config/
βββ modules/
β βββ vault-config/ # all resources (KV, PKI, policies, k8s auth, fleet, LLM)
β βββ main.tf # core: audit, kv mount, PKI, policies, k8s auth backend
β βββ ai-agents-secrets.tf
β βββ llm-secrets.tf
β βββ gitops-secrets.tf
β βββ keycloak-secrets.tf
β βββ database-secrets.tf
β βββ monitoring-secrets.tf
β βββ storage-secrets.tf
β βββ git-secrets.tf
β βββ dns-secrets.tf
β βββ vault-internal-secrets.tf
β βββ variables.tf
β βββ outputs.tf
β βββ versions.tf
βββ environment/
βββ prod/ β you are here
βββ main.tf # module "vault_config" { source = "../../modules/vault-config" ... }
βββ providers.tf # provider "vault" { β¦ } (JWT auth via Keycloak)
βββ variables.tf # provider-level + module-passthrough variables
βββ variables.secrets.tf
βββ locals.secrets.tf
βββ outputs.tf # re-exports module outputs
βββ moved.tf # 62 state migrations (delete after first apply)
βββ imports.tf # HCL imports for Vault-auto-created mounts
βββ terraform.auto.tfvars
βββ .secrets.auto.tfvars.example
βββ scripts/refresh-jwt.sh
βββ vault-config-terraform-deploy.sh
βββ terraform.tfstate
βββ README.md
Provider auth
JWT via Keycloak. scripts/refresh-jwt.sh fetches a fresh access token
from realm vault, client terraform-cicd, writes
.secrets.auto.tfvars. The provider's auth_login block exchanges that
JWT for a Vault token bound to policy terraform-cicd (TTL 1h, max 4h).
Every action attributable in the audit log to
service-account-terraform-cicd.
Why modules + environment
-
Re-use: drop a sister
environment/staging/to point the same module at a different Vault. The module is the source of truth; environment are just thin callers. -
Per-environment overrides: each env has its own
terraform.auto.tfvars,.secrets.auto.tfvars,terraform.tfstate. No cross-talk. -
Maintenance:
terraform planshows you exactly what this env diverges from. The module change is one commit; rolling it across envs is oneterraform applyper env.
Run flow
cd terraform/vault-config/environment/prod
# First time only
terraform init
# Every apply
./scripts/refresh-jwt.sh # writes .secrets.auto.tfvars (mode 0600)
terraform plan
terraform apply
Or the one-shot wrapper that does init β validate β refresh-jwt β plan β apply:
./vault-config-terraform-deploy.sh
Getting secrets to pods: the Vault Secrets Operator
The main Vault is the safe. The Vault Secrets Operator (VSO 1.3.0) is the person who fetches things out of it and delivers them to the right room. It renders ordinary Kubernetes Secrets, so an unmodified Helm chart reads them with envFrom and never learns Vault exists.
VSO works with three resources. VaultAuth says how a workload proves who it is. VaultStaticSecret copies a KV secret into a Kubernetes Secret and keeps it fresh. VaultDynamicSecret asks Vault to mint a short-lived credential on demand; it's ready to use, but nothing here uses it yet, and that's on the list at the end.
Here's the real MLflow example, end to end. First, how MLflow's service account proves who it is:
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultAuth
metadata: { name: mlops-mlflow, namespace: mlops }
spec:
method: kubernetes
mount: kubernetes
kubernetes:
role: mlops-mlflow
serviceAccount: mlops-mlflow
audiences: [vault]
tokenExpirationSeconds: 600
Then what to fetch and how to shape it:
{% raw %}
apiVersion: secrets.hashicorp.com/v1beta1
kind: VaultStaticSecret
metadata: { name: mlops-mlflow-db, namespace: mlops }
spec:
vaultAuthRef: mlops-mlflow
mount: secret
type: kv-v2
path: homelab/mlops/mlflow-db
refreshAfter: 1h
destination:
name: mlops-mlflow-db
create: true
overwrite: true
type: kubernetes.io/basic-auth # CloudNativePG insists on this type
transformation:
excludeRaw: true
templates:
username: { text: '{{ get .Secrets "username" }}' }
password: { text: '{{ get .Secrets "password" }}' }
uri: { text: 'postgresql://{{ get .Secrets "username" }}:{{ get .Secrets "password" }}@postgres-mlops-cnpg-cluster-rw.mlops.svc.cluster.local:5432/mlflow_backend_db' }
That type: kubernetes.io/basic-auth line has a story of its own in the CNPG companion article: without it, the database role gets created with no password at all, and nothing complains.
The numbers are what turn "single source of truth" from a slogan into something you can count. Today there are 48 VaultStaticSecrets across 19 namespaces, 36 VaultAuths, and exactly one VaultConnection. Not one of those Secrets was made by hand with kubectl create secret.
Why my setup looks different from the tutorials
- The tutorials teach the Agent Injector; I use VSO. The injector puts a sidecar in every pod and writes secrets to a file the app has to read. It's installed here (two pods), but no pod in the cluster uses it today.
- The tutorials stop at Shamir. Five key shares, three needed, and you type them in after every restart. That's fine for a tutorial and miserable at 3 a.m.
-
Anti-affinity is
required. I'd rather a Vault pod sitPendingthan put two Raft members on one node. The catch is in the "What's Next" list.
Terraform is the only writer
No human writes to Vault by hand. Every secret starts in a gitignored .secrets.auto.tfvars file on my machine, and Terraform writes it to Vault:
1. edit .secrets.auto.tfvars # gitignored, never committed
2. terraform apply # terraform/vault-config/environment/prod
3. Vault KV v2: secret/homelab/mlops/mlflow-db
4. VSO refreshes (hourly, or when nudged)
5. Secret mlops/mlops-mlflow-db
6. the pod's env, and CNPG's managed role password
Terraform logs in to Vault with a short-lived Keycloak JWT rather than a stored token, so even the writer has to prove who it is each time.
Today Terraform owns 47 secret paths, 55 policies and 65 Kubernetes auth roles. The paths are grouped by area under secret/homelab/. The biggest groups are monitoring, the AI agents and MLOps, each with somewhere between 15 and 19 path references.
The Vault UI's KV browser at secret/homelab/: one folder per area. Shoot the folder list only, never an opened secret.
Who can read what
Three layers keep each app in its own lane:
-
Paths. Each area owns its own folder under
secret/homelab/. - Policies. Each role's policy grants read on exactly the paths it needs. For MLflow's database secret it boils down to this:
path "secret/data/homelab/mlops/mlflow-db" {
capabilities = ["read"]
}
-
Roles bound to service accounts. The
mlops-mlflowrole only accepts a token from themlops-mlflowservice account inmlops.
So if someone took over MLflow's pod, they'd get MLflow's database password and nothing else. They couldn't read Keycloak's secrets or list other paths, and they couldn't reach the audit log.
The two months the safe stayed shut
Here's what actually happened.
The main Vault authenticates to the transit Vault with a token that Ansible created during the first install, with a 768-hour period. The playbook had a comment saying the token "auto-renews". It doesn't. A periodic token only resets its clock when something calls renew-self, and nothing did. When it ran out, every main Vault pod got permission denied from transit, couldn't unseal, and crash-looped.
The fix is a tiny daily CronJob, vault-transit-token-renewer, that renews the token. Daily against a 32-day period means it would have to fail about thirty times in a row before anything was at risk. That's deliberately generous, because the failure mode is the whole secret tier going dark.
That wasn't the only way the safe stayed shut:
-
Every reboot resealed the transit Vault. One Shamir share means someone has to type it in, and for months that someone was me, running an Ansible playbook. Now a small Deployment,
vault-transit-unsealer, checks the transit Vault every 15 seconds and unseals it if it finds it sealed. It reads the key from a mounted Secret, so its service account needs no permission to read Secrets at all. -
Power cuts left the transit pod stuck. When my hosts lose power, pods die without releasing their Ceph block-device locks, and on boot the transit pod sits in
Unknownholding a stale one. Once, that took all of Vault down until I force-deleted pods by hand. Astuck-pod-reaperCronJob now does what I used to do. -
Nobody was watching. There's now a PrometheusRule that alerts when the transit pod isn't ready. It uses pod readiness from kube-state-metrics rather than Vault's own metrics, because when Vault is sealed it's the worst possible thing to ask whether Vault is OK. The main Vault also has a ServiceMonitor scraping
/v1/sys/metricsnow.
The last power cut showed all of it working. The lab powered on, and the unsealer logged "cannot reach transit API" every 15 seconds while the transit pod came up. The moment transit answered, it logged vault-transit was SEALED -> unsealed OK. About two minutes later the main Vault pods were up and unsealed. I was making coffee.
kubectl -n vault-transit logs deploy/vault-transit-unsealer: the "cannot reach" lines, then SEALED -> unsealed OK. It's the most reassuring log line in the cluster.
How reliable is it now?
Here's what the safe survives, and the evidence for each:
| What goes wrong | What happens | How I know |
|---|---|---|
| One Vault pod dies | The two standbys keep serving, and Raft elects a new leader if it was the active one | Three-node Raft, HA Mode: active plus two standbys |
| A node reboots | The pod comes back, auto-unseals against transit and rejoins Raft | After the last power cut, vault-1 and vault-2 came back and rejoined on their own, with no crash restarts |
| The whole lab loses power | Transit is unsealed within about 15 seconds of coming up, and the main Vault follows | Last power cut: transit unsealed the moment it answered, main Vault up about two minutes later, no human involved |
| The auto-unseal token ages | Renewed daily against a 32-day period |
vault-transit-token-renewer runs every day |
| A pod is stuck after a power cut | Reaped within three minutes |
stuck-pod-reaper runs every 3 minutes |
| The transit Vault isn't ready | An alert fires | PrometheusRule on transit readiness |
| Vault is down for a while | Running apps keep their existing Secrets; only new syncs wait | That's how the two-month outage went unnoticed |
And the honest limits, all on the list at the end: the transit Vault is a single pod, the main Vault spreads across nodes rather than strictly across datacenters, and there are no scheduled Raft snapshots yet.
The traps that taught me
Istio makes a TLS listener look like plain HTTP
The pods came up green, then vault status said http: server gave HTTP response to HTTPS client. The listener was TLS. Istio's sidecar was grabbing the connection and forwarding plain text into it. Exclude Vault's ports in both directions:
server:
annotations:
traffic.sidecar.istio.io/excludeInboundPorts: "8200,8201"
traffic.sidecar.istio.io/excludeOutboundPorts: "8200,8201"
Everyone remembers inbound. Forgetting outbound breaks retry_join, because vault-1 calling vault-0 goes through the sidecar too. And the chart key is server.annotations, not server.podAnnotations, which the chart silently ignores.
Raft redirects to an IP your certificate doesn't cover
vault-1 and vault-2 joined, got redirected to the leader's pod IP, and died with x509: cannot validate certificate for 10.42.x.y because it doesn't contain any IP SANs. Pod IPs change on every reschedule, so they don't belong in a certificate. Pin the name instead:
retry_join {
leader_api_addr = "https://vault-0.vault-internal:8200"
leader_tls_servername = "vault.vault.svc.cluster.local"
}
One letter in the seal block
Every pod crash-looped with seal init failed and nothing else. The transit Vault runs without TLS inside the cluster, so its address has to be http://. Every example online uses https://. I stared at that line for an hour before I saw it.
vault operator init fails one time in three
A fresh vault-0 spends its first few seconds in a busy retry_join loop, and init during that window returns a bare 500 internal error. Ansible's no_log hid even that. The fix is to wait until vault status settles, pause, then init with retries, and treat "already initialized" as success.
The unseal command that ignores your key
Older notes of mine, and the first version of this article, said to pipe the key in with vault operator unseal -. On Vault 1.21.2 that doesn't read stdin. It treats the - as the key itself, and the unseal quietly fails. The unsealer passes the key as an argument instead. If you're doing it by hand, read it into a variable first so it never lands in your shell history:
read -rs KEY
kubectl -n vault-transit exec vault-transit-0 -- vault operator unseal "$KEY"
unset KEY
The operator says 38 secrets are broken. They aren't.
After the last power-on, kubectl get vaultstaticsecrets -A showed 38 of 48 with SecretSynced=False. The message was a no route to host to an old pod IP, logged while Vault was still coming up. Every one of those 38 Secrets exists and has its keys. The condition just hadn't been updated since. Check the Secret itself, not the condition.
What I thought Vault was doing for my disks
An earlier draft of this article said Vault holds the encryption keys for my Ceph disks. It doesn't. The Ceph config points at Vault, a rook mount and a policy exist, and the disks really are encrypted. But the keys are in Ceph's own monitor store. Vault's rook mount is empty.
The cause is the Istio pattern again. A DestinationRule makes the sidecar wrap traffic to Vault in TLS, so meshed clients have to speak http://. Rook was set to https://, which is TLS inside TLS, and it never once connected:
TLSv1.3 (OUT), TLS alert, record overflow (534)
error:0A0000C6:SSL routines::packet length too long
Two things came out of finding that. The VAULT_SKIP_VERIFY setting I'd been apologising for never mattered. The thing actually skipping certificate checks is insecureSkipVerify: true in that DestinationRule, and it applies to every meshed client of Vault. And my storage doesn't depend on Vault at all, which is a happier failure mode than the one I thought I had. I wrote that Vault held the keys because the config said so. Config isn't evidence; ceph config-key ls is.
The CephCluster spec showing KMS_PROVIDER: vault, next to ceph config-key ls | grep dm-crypt listing six dm-crypt/osd/<fsid>/luks entries. Shoot the spec and the key names only, never a key, and never the rook/ KV path.
What it costs
- Pods: 3 main Vault, 1 transit, 1 unsealer, 2 Agent Injector, the VSO controller, and a CSI provider on every node (29 on my cluster).
-
Storage: 3 Γ 10Gi Raft plus 3 Γ 5Gi audit on
ceph-block, and 1Gi for transit. - Upkeep: the daily renewer, the reaper, and an alert. All three exist because of something that went wrong.
What it buys: one answer to "where does this password come from?", an audit log of every read, rotation by terraform apply, and per-app isolation enforced by Vault rather than by everyone being careful.
What's Next?
The safe now opens itself after a power cut, and I'll know if it doesn't. What's still on my list, roughly in order of how bad the failure would be:
-
Spread Vault by datacenter, not just by node. The anti-affinity uses
kubernetes.io/hostname. Today the pods happen to sit one per datacenter, but nothing stops two landing in the same one. If that datacenter goes, Raft loses quorum and the safe stays shut. The CNPG clusters already do this properly withtopology.kubernetes.io/zone. -
Scheduled Raft snapshots. Ceph's three copies aren't a backup.
vault operator raft snapshot saveon a CronJob is. -
A NetworkPolicy for the
vaultnamespace. The transit Vault has one. The main Vault doesn't, so anything in the cluster can reach port 8200 and try its luck. -
Close
insecureSkipVerifyby getting Vault's CA to the sidecars, then point Rook athttp://so disk keys can finally reach Vault. - Dynamic database credentials, so apps get short-lived Postgres logins instead of static passwords.
Next episode, the secrets get put to work. The NVIDIA GPU Operator brings the cluster's GPUs online, and everything that trains a model starts asking the safe for its keys.
If you're still keeping secrets in a .env you swear is gitignored, this is the episode to borrow from. Just set an alarm for your unseal token.
Author: George Ezejiofor
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.













