ragleap-ops · release notes
2026-09-06 · ragleap-ops-v0.3.0
ragleap-ops v0.3.0
Added
- Backup/DR:
backup-pvc.yaml, db-backup-cronjob.yaml (daily pg_dump, no downtime), neo4j-backup-cronjob.yaml (scale-to-zero + neo4j-admin dump, since Community Edition has no online backup command).
Verified
pg_dump live-tested against a real running database, produced a valid dump.
- neo4j scale-to-zero + dump pattern live-tested end-to-end via a real Job, produced a genuine 257.8MiB/36-file dump successfully.
- CronJob wrapper (RBAC, scheduling) validated structurally, not exercised via an actual cron trigger this session.
Known limitations
- Neo4j backup is fail-loud by design: a failed dump leaves neo4j scaled to 0 rather than auto-restoring, to avoid a silently-failing backup going unnoticed. Requires monitoring CronJob/Job status separately.
Fixed
k8s/neo4j-deployment.yaml's liveness probe still had the original, pre-fix timing (initialDelaySeconds: 20, no explicit failureThreshold) even though the equivalent bug was already found and fixed in the Helm chart's neo4j template earlier this session. The two deployment paths had silently diverged. Found while live-testing backup/DR tooling on a fresh cluster -- neo4j genuinely crash-looped (kubelet killing it ~7 seconds into JVM startup, well before the database could bind its HTTP listener). Fixed to match the Helm chart's already-proven values (initialDelaySeconds: 60, failureThreshold: 6), then re-verified stable (0 further restarts after the one within the expected startup window).
Added
- Environment overlay files (
values-dev.yaml, values-staging.yaml, values-prod.yaml) for the Helm chart, layered on top of the base values.yaml.
app.replicaCount/voice.replicaCount parameterized (were previously hardcoded to 1 in the templates, which would have silently made environment-based replica scaling impossible). db/neo4j intentionally remain hardcoded at 1 replica -- both are stateful, single-writer services on ReadWriteOnce PVCs.
Verified
- Rendered output confirmed to differ correctly across all three overlays via
helm template -f values-<env>.yaml: dev (1/1 replicas, ingress disabled), staging (1/1 replicas, ingress enabled with its own hostname), prod (3/2 replicas, ingress enabled with its own hostname).
Added
- Optional Helm chart Ingress + cert-manager Certificate (
ingress.enabled, default false) for exposing the app outside the cluster over TLS.
Verified
- Full Ingress + TLS chain live-tested end-to-end on a real kind cluster with a genuine NGINX Ingress Controller and cert-manager, using a self-signed ClusterIssuer. Confirmed via openssl that the correct certificate (matching SNI hostname) was served, not a generic fallback. HTTP routing through the Ingress to the backend independently confirmed.
Known limitations
- Production Let's Encrypt issuance was not live-tested this session — real ACME HTTP-01 challenges require public DNS and an internet-reachable port 80, which a local kind cluster cannot satisfy. The provided ClusterIssuer example is the standard, documented cert-manager pattern, not independently verified against a real domain.
Added
- NetworkPolicy resources (both
k8s/ and helm/ragleap-ops/) restricting ragleap-db ingress to ragleap-app/ragleap-voice only, and ragleap-neo4j ingress to ragleap-app only (ragleap-voice doesn't use neo4j in the current codebase, confirmed by checking real source, not assumed).
Verified
- NetworkPolicy enforcement mechanism live-tested end-to-end on a real kind + Calico cluster: unlabeled traffic genuinely blocked (timeout, exit code 1), correctly-labeled traffic genuinely allowed (exit code 0) — not just applied without error.
- Real label selectors cross-checked against actual Deployment manifests — confirmed exact match.
Known limitations
- kind's default CNI does not enforce NetworkPolicy at all; testing requires Calico or another NetworkPolicy-capable CNI (documented in README).
- Full ragleap stack (db+app+voice+neo4j) was not live-tested together under Calico in this session due to genuine VPS memory constraints — Calico's own baseline overhead left insufficient headroom for reliable 4-service testing on this specific host. Enforcement mechanism and label correctness were verified separately, not as one combined integration test.
Release notes mirrored from GitHub Releases.