Kubernetes deployment manifests for RagLeap Core.
Ships Deployment + Service manifests for all four docker-compose.yml
services (db, app, voice, neo4j), translated from the real
compose file and live-tested end-to-end on a kind cluster — all four
pods reached 1/1 Running with zero restarts after a probe-timing bug
was found and fixed via live testing.
Contents
k8s/— raw Deployment, Service, PVC, Secret, and ConfigMap manifests- Apply order: Secrets/ConfigMap → PVCs →
db→ everything else (app/voicedepend ondbvia an init container that pollspg_isready)
Regenerating local-only files
Two files are intentionally not committed (see .gitignore) since
they'd otherwise bake real secret values into git history:
kubectl create secret generic ragleap-app-env \ --from-env-file=$HOME/ragleap-core/.env \ --dry-run=client -o yaml > k8s/app-env-secret.yaml
Secrets management for production (Sealed Secrets)
The kubectl create secret --from-env-file pattern above is fine for local testing, but it leaves real credentials sitting as plaintext files and shell history. For production, use Sealed Secrets instead (https://github.com/bitnami-labs/sealed-secrets) -- secrets get encrypted against your specific cluster's public key, so the encrypted result is safe to commit to git. No other cluster can decrypt it.
One-time setup (per cluster):
kubectl apply -f https://github.com/bitnami-labs/sealed-secrets/releases/download/v0.39.1/controller.yaml
Sealing a secret (works offline once you have exported your cluster's public cert -- no live cluster connection needed at seal-time, safe to run in CI):
kubeseal --controller-name sealed-secrets-controller \ --controller-namespace kube-system --fetch-cert > sealed-secrets-pub-cert.pem kubectl create secret generic ragleap-db-secret \ --dry-run=client \ --from-literal=POSTGRES_USER=ragleap \ --from-literal=POSTGRES_PASSWORD=your-real-password \ --from-literal=POSTGRES_DB=ragleap_core \ -o yaml | kubeseal --format yaml --cert sealed-secrets-pub-cert.pem > db-secret-sealed.yaml kubectl apply -f db-secret-sealed.yaml
Important: a sealed secret is tied to the exact cluster whose public key encrypted it. A SealedSecret sealed for one cluster will not decrypt on a different cluster -- never copy a sealed file between environments; reseal against each target cluster's own cert instead.
Status
Both the raw manifests (v0.1.0) and the Helm chart (v0.2.0, helm/ragleap-ops/) are live-verified end-to-end on a real kind cluster — see CHANGELOG.md for the specific bugs found and fixed.
NetworkPolicies (restricting pod-to-pod traffic)
Both k8s/ and helm/ragleap-ops/ include NetworkPolicy resources restricting which pods can reach ragleap-db (port 5432, from ragleap-app and ragleap-voice only) and ragleap-neo4j (ports 7474/7687, from ragleap-app only -- ragleap-voice does not use neo4j in the current codebase).
Important: kind's default CNI (kindnet) does not enforce NetworkPolicy at all -- policies will apply without error but traffic will not actually be blocked. Testing NetworkPolicy enforcement requires a CNI that supports it, such as Calico:
kind create cluster --config kind-calico-config.yaml # disableDefaultCNI: true kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/v3.28.0/manifests/tigera-operator.yaml kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/v3.28.0/manifests/custom-resources.yaml
Verification performed: the enforcement mechanism itself was live-tested end to end on a real kind + Calico cluster (a labeled pod was correctly blocked from an unlabeled target, then allowed once correctly labeled -- confirmed via real timeout/success exit codes, not assumed). The real ragleap-db/ ragleap-neo4j label selectors were verified to match the actual Deployment labels used elsewhere in this chart. A full live run of the complete ragleap stack under Calico was not completed in this session due to genuine memory constraints on the test VPS (which also runs unrelated production services) -- Calico's own baseline overhead left too little headroom for a 4-service stack reliably. This is an honest scope boundary, not a claim that the policies were fully integration-tested against the live application.
Ingress + TLS
The Helm chart includes an optional Ingress + cert-manager Certificate
(disabled by default via ingress.enabled: false). Enabling it requires:
- An NGINX Ingress Controller installed on your cluster
- cert-manager installed, with a real ClusterIssuer configured (Let's Encrypt for production, or a self-signed issuer for local testing)
- A real hostname you own, set via
ingress.hostnamein values.yaml - Your ClusterIssuer's name set via
ingress.clusterIssuer
Example production ClusterIssuer (Let's Encrypt, replace the email):
cat <<EOF | kubectl apply -f -
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: [email protected]
privateKeySecretRef:
name: letsencrypt-prod-key
solvers:
- http01:
ingress:
ingressClassName: nginx
EOF
Then deploy with:
helm install ragleap-ops ./ragleap-ops \ --set ingress.enabled=true \ --set ingress.hostname=your-real-domain.com \ --set ingress.clusterIssuer=letsencrypt-prod
Verification performed: the full Ingress + TLS chain was live-tested end to end on a real kind cluster with a genuine NGINX Ingress Controller and cert-manager, using a self-signed ClusterIssuer (Let's Encrypt itself requires public DNS and internet-reachable ports, which a local kind cluster cannot satisfy for a real ACME challenge). Confirmed via openssl: the correct certificate was served over TLS via SNI, matching the exact hostname requested, not a generic fallback certificate. HTTP routing through the Ingress to the backend service was independently confirmed working. Production Let's Encrypt issuance was not live-tested in this session for the reason above -- the issuer configuration shown is the standard, documented cert-manager pattern, not independently verified against a real public domain here.
Multi-environment deployment
Environment-specific overrides live in helm/ragleap-ops/values-{dev,staging,prod}.yaml,
layered on top of the base values.yaml (never edit the base file for a
single environment):
helm install ragleap-ops ./ragleap-ops -f values-dev.yaml # local/dev helm install ragleap-ops ./ragleap-ops -f values-staging.yaml # staging helm install ragleap-ops ./ragleap-ops -f values-prod.yaml # production
db and neo4j intentionally have no replicaCount in any environment
file and stay hardcoded at 1 in the templates -- both are stateful,
single-writer services backed by ReadWriteOnce PVCs. Horizontal scaling
for either would require a genuinely different storage/clustering
architecture, not a values.yaml change, so this isn't offered as a
configurable option that would silently produce a broken multi-writer
setup if someone bumped the number.
Verified: rendered output for app.replicaCount/voice.replicaCount
and ingress.enabled/ingress.hostname confirmed to differ correctly
across all three overlay files via helm template -f values-<env>.yaml
-- dev (1/1, no ingress), staging (1/1, ingress enabled with its own
hostname), prod (3/2, ingress enabled with its own hostname).
Backup / Disaster Recovery
Scheduled backups for both db (Postgres) and neo4j live in
k8s/backup-pvc.yaml, k8s/db-backup-cronjob.yaml, and
k8s/neo4j-backup-cronjob.yaml.
Postgres: pg_dump runs against the live database daily -- no
downtime, Postgres handles concurrent dumps natively via MVCC.
Neo4j: Community Edition has no online/hot backup command (neo4j-admin database backup is Enterprise-only, confirmed via --help against the
real image -- only dump exists in Community). The CronJob scales
ragleap-neo4j to 0 replicas, waits for the pod to fully terminate, then
runs neo4j-admin database dump against the now-unlocked PVC via a Job
with the same volume mounted.
Important design decision -- fail loud, no automatic restart: if the
dump step fails, neo4j is intentionally left scaled to 0 rather than
automatically restored. Coupling application uptime to backup success
was considered and rejected -- a silently-failing backup with automatic
scale-up could leave a real backup gap unnoticed for a long time while
the app appears healthy. Monitor CronJob/Job success explicitly
(kubectl get cronjobs, kubectl get jobs, or your alerting stack) and
manually restore service after confirming/investigating a failure:
kubectl scale deployment ragleap-neo4j --replicas=1
Verification performed: both pg_dump and the scale-to-zero +
neo4j-admin dump pattern were live-tested against real running
instances on a real kind cluster -- pg_dump produced a valid dump
against the live database; the neo4j pattern produced a real, complete
257.8MiB dump (36 files) via the Job mechanism. The scheduled CronJob
wrapper itself (RBAC, scale-down init container, timing) was validated
structurally (python3 -c "import yaml..." parse + document/kind
checks) but not exercised end-to-end via an actual cron trigger this
session.