Metrics, logs, and alerting for RagLeap Core -- Prometheus, Grafana, Loki, and AlertManager, wired against the real db/app/voice/neo4j services.
Status
Prometheus + postgres_exporter, Grafana, Loki + Promtail, and
AlertManager are all built and live-verified end-to-end against a
real cluster:
postgres_exporterconnects toragleap-dbusing a read-only monitoring role;/metricsreturns real, non-zero PostgreSQL statistics (pg_up 1, livepg_stat_databasevalues).- Grafana's Prometheus and Loki datasources are both confirmed live through Grafana's own datasource proxy, not just "pod is Running."
- Loki + Promtail log shipping is proven end-to-end: 21+ real log
streams confirmed in Loki with correct
namespace/pod/containerlabels, spanning nearly every pod in the cluster. See CHANGELOG for the full diagnostic history -- several real, non-obvious bugs (a cross-namespace DNS lookup, a Kubernetes service-discovery__path__resolution issue, and a YAML document-boundary corruption) were found and fixed to get here. - AlertManager is wired end-to-end to Prometheus and live-verified:
Prometheus's own
/api/v1/alertmanagersconfirms it as a genuinely registered target, and a real first alert rule (PostgresExporterDown) loads correctly. No real notification receiver is configured yet -- routing, grouping, and silencing all work for real, but nothing will actually notify a human until a real Slack/email/PagerDuty receiver replaces the placeholder webhook invalues.yaml.
SLO/SLI dashboards are not yet built -- correctly sequenced after alerting has proven reliable, per the build order below.
Planned build order
- Prometheus + exporters (postgres_exporter for db) -- done, live-verified. neo4j's exporter remains disabled pending independent verification, see "Known open items".
- Grafana, provisioned dashboards-as-code, verified against real flowing data -- done, live-verified.
- Loki + log shipping -- done, live-verified end-to-end.
- AlertManager -- done, wired and live-verified. No real receiver configured yet; see "Known open items".
- SLO/SLI dashboards, only after real data has accumulated for days, not minutes
Design principles
- RagLeap-specific scoping, same honest choice
ragleap-opsmade for itself -- not a generic monitoring tool - Every claim live-verified against a real cluster or explicitly labeled unverified, same discipline as every other package here
Known open items
- Neo4j Prometheus support is unverified and conflicting in the wild.
Neo4j's own KB and a working xk6-neo4j example show
metrics.prometheus.enabled=trueconfigured successfully against Neo4j Community 4.4.x. A separate monitoring vendor's own compatibility notes claim Community Edition is unsupported for their specific collector.neo4jExporter.enableddefaults tofalseinvalues.yamluntil this is live-verified against the realneo4j:5-communityimage in this repo's ownkindcluster. - Loki's retention (
168h/ 7 days) is unverified for real storage-sizing needs; flagged invalues.yamlas a placeholder. - AlertManager has no real notification receiver configured. Alerts
are correctly routed and grouped but not delivered anywhere. Replace
the placeholder webhook in
values.yamlwith a real Slack/email/ PagerDuty receiver before relying on this for real incident response. - AlertManager's state (silences, notification log) uses
emptyDir, not a PVC -- lost on pod restart. - A recurring log-shipping health check remains unbuilt.
- No canary or blue-green deployment strategy exists in this project
-- every Deployment uses the Kubernetes default
RollingUpdateor, for Prometheus specifically, an explicitRecreate(required by its single-writer TSDB storage). Correctly sequenced after GitOps tooling (ArgoCD/Flux), which has not been started.