Added
- AlertManager deployment (Deployment, Service, ConfigMap), wired end-to-end to Prometheus.
- First real alert rule:
PostgresExporterDown, targeting a live, proven service. - New
ragleap-prometheus-rulesConfigMap mounted into Prometheus.
Fixed
- Prometheus's Deployment defaulted to RollingUpdate, but its ReadWriteOnce PVC holds a single-writer TSDB directory. The first real rollout crash-looped with
lock DB directory: resource temporarily unavailable. Latent since the original deployment, exposed by this work. Fixed withstrategy: Recreate.
Verified
- Prometheus
/api/v1/rulesconfirms the rule loaded (health ok);/api/v1/alertmanagersconfirms AlertManager registered as an active target; AlertManager/-/readyreturns OK.
Known limitations
- No real notification receiver configured. Alerts are routed and grouped but not delivered to any human until a real Slack/email/PagerDuty receiver replaces the placeholder webhook.
- AlertManager state uses
emptyDir(lost on restart). - No canary or blue-green deployment strategy exists in the project; correctly sequenced after GitOps tooling (step 14).
Full details: CHANGELOG.md