Added
- Prometheus server and Grafana, live-verified end-to-end. Deployment/Service/PVC templates for Prometheus, plus Grafana (Deployment, Service, Secret, datasource ConfigMap).
- Loki deployment and Promtail DaemonSet for log shipping, live-verified end-to-end: 21+ real log streams confirmed with correct namespace/pod/container labels.
- Grafana datasource ConfigMap conditionally wires in Loki alongside Prometheus.
- New
loki:andpromtail:blocks in values.yaml.
Fixed
postgres-exporterrunAsNonRoot fix (runAsUser: 65534) for named-user images.- Cross-namespace DNS lookup failure in Promtail's Loki client — fixed by using the fully-qualified service name.
kubernetes_sd_configsproduced 0 active targets: fixed a node-scoping HOSTNAME issue and replaced a custom regex with Grafana's canonical__path__pattern, plus added thecripipeline stage.- A missing newline had silently corrupted a ConfigMap's YAML document boundary, causing a cascade of confusing symptoms across several
helm upgradecycles. - Raised Loki's ingestion rate limit to handle replay bursts from Promtail restarts.
- Removed the temporary
static-pod-logsdiagnostic job, which was silently starving the real job of every file via shared position tracking.
Verified
- Full pipeline confirmed live, layer by layer (Prometheus targets, Grafana health/datasources, Loki series query returning 21 real labeled streams, Promtail
/readyreturningReady).
Known limitations
- Neo4j Prometheus support unverified;
neo4jExporterstays disabled. - AlertManager, SLO/SLI dashboards, recurring log-shipping health check remain unbuilt.
- Loki retention (7 days) unverified for real sizing needs.
Full details: CHANGELOG.md