Observability

Prometheus Setup

Bloodraven exposes Prometheus metrics on the operator metrics Service, port 8080 named metrics.

Helm values

Deployment coordination adds active/age gauges and grant, expiry, revocation, and planned-deferral counters. Their exact names and labels are in the Deployment API metric reference. Keep revocation and fencing alerts active during application maintenance; a hold is not permission to suppress an emergency failover alert.

metrics:
  service:
    enabled: true
  serviceMonitor:
    enabled: true
    interval: 30s
    scrapeTimeout: 10s
    labels:
      release: kube-prometheus-stack
helm upgrade --install bloodraven bloodraven/bloodraven \
  --namespace bloodraven \
  --create-namespace \
  --values bloodraven-values.yaml

ServiceMonitor

The chart renders a ServiceMonitor when metrics.serviceMonitor.enabled=true. It selects the operator metrics Service in the release namespace.

If you manage the monitor yourself:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: bloodraven
  namespace: bloodraven
  labels:
    release: kube-prometheus-stack
spec:
  namespaceSelector:
    matchNames:
      - bloodraven
  selector:
    matchLabels:
      app.kubernetes.io/name: bloodraven
  endpoints:
    - port: metrics
      interval: 30s
      scrapeTimeout: 10s

Plain Prometheus scrape config

scrape_configs:
  - job_name: bloodraven
    kubernetes_sd_configs:
      - role: endpoints
        namespaces:
          names:
            - bloodraven
    relabel_configs:
      - source_labels: [__meta_kubernetes_service_label_app_kubernetes_io_name]
        action: keep
        regex: bloodraven
      - source_labels: [__meta_kubernetes_endpoint_port_name]
        action: keep
        regex: metrics

Verify targets

kubectl get service -n bloodraven
kubectl get servicemonitor -n bloodraven
kubectl port-forward -n bloodraven deploy/bloodraven 8080:8080
curl http://localhost:8080/metrics | grep '^bloodraven_'

In Prometheus, check Status > Targets for the bloodraven job or ServiceMonitor-generated target.

Reader and source-convergence monitoring

bloodraven_replication_source_state{namespace,group,site,state} is a state-set gauge for every follower. The namespace and group labels keep identically named sites in different failover groups in separate series. It emits the bounded state values converged, pending, and blocked; exactly one is 1 for a follower and the others are 0. Active primaries have no active source state. Combine it with bloodraven_replication_running{namespace,group,site,role,thread} and bloodraven_replication_lag_seconds{namespace,group,site,role} when alerting on a reader.

Reader failures are deliberately isolated from the failover group's shared Ready and Degraded conditions. A reader can be unreachable, lagging, or source-blocked while the core candidate/DR topology remains Ready and not Degraded. Alert on reader sites by role rather than by site name. For core RPO (promotable and DR replicas, not designed-to-lag readers):

bloodraven_replication_lag_seconds{role=~"primary-candidate|dr-only"} > 30

The issue-142 form that pages only promotable replicas is:

bloodraven_replication_lag_seconds{role="primary-candidate"} > 30

Alert on a specific reader by namespace, group, and site:

bloodraven_replication_source_state{namespace="warehouse",group="orders",site="reader",state!="converged"} == 1
bloodraven_replication_running{namespace="warehouse",group="orders",site="reader"} == 0
or bloodraven_replication_lag_seconds{namespace="warehouse",group="orders",site="reader"} > 30

To page every role: read-only site under the operator, drop the identity labels and keep {role="read-only"}.

Use a sustained for interval appropriate to the normal poll and recovery cadence. A blocked source commonly needs GTID investigation and possibly a cold reclone; a pending source may clear on a later bounded retry.

License update-period reminder

This is a compliance reminder, not an outage. The operator does not change behavior when the update period ends.

(bloodraven_license_updates_expiry_timestamp_seconds - time()) / 86400 < 30

Alerts

Keep alert rules with your platform monitoring stack. Alert names should link to Alert To Runbook Map, and metric details live in Monitoring Reference.