Kubernetes operator for MySQL

Give your MySQL topology a survival instinct.

Bloodraven runs MySQL async replication failover groups across sites. It owns detection, fencing, promotion, DNS steering, clone bootstrap and backups — so losing an entire site is a status change, not an incident.

live topology mysqlfailovergroup/orders poll 2s
iad → pdx · async
PRIMARYIADus-east-1a · 10.0.1.1read_only=OFF
REPLICAPDXus-west-2a · 10.0.2.1lag 0s · source iad
READ-ONLYreaderlag 0s · serving readssource iad
fencing armed · exactly one writer
ACTIVE DNS · TTL 60orders.az.example.com10.0.1.1iad

operator ready · watching 3 sites · press Kill the active site to watch a failover.

T+00.0s

Scripted replay of the documented promotion sequence. Kill the site and read the clock.

~6s
to declare a site unreachable (3 failed 2s polls)
RPO 0
on a planned switchover, by construction
47+
scripted chaos scenarios run in CI
1
CRD, one controller, one reconcile loop

Sixty pages of docs, one prompt

Don't read 60 pages. Just ask.

The assistant is grounded in this entire documentation set — the CRD reference, every runbook, the failure-mode matrix and the log schema. Ask it in your own words and it answers with links to the exact page.

Bloodraven AIGrounded in this documentation setready

Start here

Day-2+ Diagnostic Copilot

Meet Bloodraven Doctor. AI diagnostics in the CLI.

Give your AI coding assistant deep diagnostic vision over live MySQLFailoverGroup clusters. bloodraven-doctor performs 60-second non-destructive triage, detects GTID divergence, audits keyring status, and crafts safe step-by-step remediation plans.

Install for any AI agentCursor · Claude Code · Antigravity · Windsurf · Roo
$npx skills add shipstream/bloodraven
60-Second Fast Triage
Non-destructive assessment of quorum, active primaries, split-brain status, and replication lag.
GTID Divergence Auditor
Compares executed transaction sets across sites to detect errant writes and guide safe recloning.
Keyring Escrow Safety
Enforces strict safety guards to protect master encryption keys from loss during pod restarts.
Runbook-Grounded Plans
Produces step-by-step remediation commands with clear blast radius and RPO trade-offs.
bloodraven-doctor · orders-productionactive
agent>Run bloodraven-doctor on namespace orders
🩺 Bloodraven DoctorInspecting MySQLFailoverGroup: orders
Status: 🟡 DEGRADED (Replication Lag Detected)
Active Primary: iad (Writable, super_read_only=OFF)
Standby Site: pdx (Replica, lag=48s, IO_Thread=ON)
Keyring Status: Sealed (Digest: 8f3a... matched escrow)
Dragonfly: Ready (replTakeoverSupported=true)
Root Cause Analysis:
High replication lag on standby `pdx` due to large transaction batch on primary. GTIDs are sequential; no errant transactions detected.
Recommended Safe Remediation:
1. Monitor catchup: kubectl bloodraven status orders -n orders2. Preflight failover: kubectl bloodraven promote orders --to=pdx (blocked until lag=0s)
Blast Radius: Zero (Read-only standby catching up) · Data Loss Risk: None

Bootstrap, scale out, encrypt, provision, back up

Data that looks after itself.

Bootstrap, read scale-out, encryption, tenant provisioning and backup are part of the operator, not five more systems you have to wire together.

CLONE

Clone-based bootstrap

New replicas seed themselves with MySQL's clone plugin and pick up replication with GTID auto-positioning. No mysqldump window, no snapshot juggling, no manual data transfer. The same path repairs a site whose PVC was lost, or one you deliberately reclone after divergence.

CLONE INSTANCE FROM 'replicator'@'orders-iad:3306' …
CHANGE REPLICATION SOURCE TO SOURCE_AUTO_POSITION = 1;
START REPLICA;
Explore
AES

Data-at-rest encryption, for free

InnoDB tablespace encryption on ordinary PVCs, using the GPL keyring component that ships with MySQL Community Edition. No Oracle Enterprise licence, no encrypted CSI storage class, and the master key never lands on the data PVC or a worker-node disk. Rotation, sealing and escrow are handled by the operator.

spec:
  encryptionAtRest:
    enabled: true
Explore
RO

Read-only replicas

Append a read-only site and it follows whichever site is active — never promoted, never a planned-failover target, never a DNS target, never a clone donor. It gets its own client Service and its own mysqlConf, so you can size it for reporting, analytics or a CDC tap without touching the write path. Fall behind readOnlyMaxLagSeconds and its endpoint sheds until it catches up; lose its data and it reclones itself.

- name: reader
  role: read-only
  zone: us-east-1b
Explore
CRD

Tenant databases as resources

Declare a schema, its owner and any extra principals — a per-tenant SELECT-only support reader, say — as a MysqlDatabase, and the operator runs the SQL on whichever site is primary. Every account is Secret-backed, scoped to its own schema and, if you like, to a list of source IPs; rotation is a Secret write, and removing an entry revokes and drops it. Nothing but Bloodraven holds the admin credential.

users:
  - secretName: support-ro-mysql
    privileges: [SELECT]
    hosts: [35.1.2.3, 35.4.5.6]
    resourceLimits:
      maxUserConnections: 5
Explore
B/R

Backup and restore. The whole nine yards.

Not a cron job that shells out to mysqldump. Backup is a first-class part of the operator, from the schedule all the way through to proving the artifact can actually be restored.

  • S3 or PVC — object storage for durability, PVC for labs
  • Scheduled and on-demand — a MysqlBackup whenever you need one
  • Structured retention — plus exponential-backoff retries on failure
  • Point-in-time recovery — binlog archiving between full dumps
  • Encrypted artifacts — application-level, independent of the store
  • Prometheus metrics — every run, every failure, every sweep
  • Automatic cleanup — deleting the object removes its artifacts
  • Verification — load the dump into a throwaway MySQL and prove it
  • initFromBackup — bootstrap a brand-new failover group from a dump
  1. 01Dumpconsistent, scheduled or on demand
  2. 02RetainS3 or PVC · structured retention
  3. 03Verifyrestore into a throwaway MySQL
  4. 04BootstrapinitFromBackup into a new group

No quorum. No coordinator.

One controller. One CRD. One loop.

There is no distributed consensus in Bloodraven, no coordinator to keep quorum, and no second system to reconcile against. A single reconcile loop reads the observed topology and writes the decision — which is why the failure modes fit on one page.

No coordination problem
Nothing to elect, nothing to split. State lives in the CR's status and in MySQL itself.
The data plane doesn't need the operator
A healthy primary and replica keep serving reads and writes with zero operator involvement — it is on the detection path, not the request path.
Correct even while the operator is down
Sidecars self-fence when they can reach neither the operator nor their peer, so no split brain is possible while the control plane is missing.
Every failure mode is written down
A documented matrix of faults, what the operator does, how long it takes, and what it costs you.
Read the failure-mode matrix
orders.yaml
 1apiVersion: shipstream.io/v1alpha1 2kind: MysqlFailoverGroup 3metadata: 4  name: orders 5spec: 6  dns: 7    hostname: orders.az.example.com 8    ttl: 60 9  replication:10    readOnlyMaxLagSeconds: 3011  sites:12    - name: iad13      zone: us-east-1a14      lbIP: 10.0.1.115    - name: pdx16      zone: us-west-2a17      lbIP: 10.0.2.118    - name: reader19      role: read-only20      zone: us-east-1b

Deployment API · /deploy/v1

Migrations that know the ground moved.

A deploy pipeline asks Bloodraven for a lease before it touches the schema. The lease serializes migrations across the group, holds planned failovers still while it is renewed, and is revoked the instant the topology generation changes — so a migration can never keep running on a primary it did not start on.

lease lanes https://bloodraven:8443/deploy/v1/groups/ordersttl 15s · renew every 5s

Two deploys collide. One migration per group. The second deploy gets 409 held, with the holder, and waits its turn.

a1c3…deploy · orders-appIDLEb7f0…deploy · orders-appIDLEplanned opsNONEtopologygen 17migrationfailover-holdmigrationfailover-holdT+05s10s15sT+00.0s

deploy client ready · two deploys collide · press Play to watch the API arbitrate.

T+00.0s
MUTEX One migration per group
A second deploy gets 409 held with the holder's operation ID and expiry. Two pipelines cannot race the same schema, and a crashed one stops blocking when its TTL lapses.
HOLD Planned disruption waits
A renewed failover-hold defers planned promotion, ordered updates and restore-in-place with phase Deferred and reason DeploymentHold. It never delays emergency failover or primary fencing.
FENCE Revoked, not reconnected
Every lease is stamped with topologyGeneration. Any authoritative site change bumps it and the next renewal returns 409 revoked, so the client stops instead of continuing DDL against the new primary.
RBAC No cluster-admin for CI
Callers authenticate with a projected ServiceAccount token and are authorized from MysqlDatabase.spec.deploymentClients. No permission to patch the group, no Secrets, no MySQL credentials, TLS only.
Read the deployment API reference

Train the on-call rotation

Bloodraven in Production.

Train your on-call team on a real cluster, not on slides. Hold a site down and time the promotion. Audit the exact transactions an emergency failover cost you. Wire an application that survives one. Roll a MySQL upgrade underneath live traffic.

7
units
28
topics
35
quizzes and tests

Every number in the course traces to a source: operator code, the shipped CRDs, recorded chaos-run forensics, and the MySQL manual.

  1. 1Meet the group — stand up a real three-site failover group and read its status.
  2. 2How the operator decides — predict its next move from a status dump alone.
  3. 3Emergency failover end to end — time the promotion, audit what it cost.
  4. 4Where failover meets your application — pools, reconnects and stale reads.
  5. 5When the world misbehaves — self-fencing, split brain, five kinds of partition.
  6. 6Backups, disaster recovery and a go-live checklist you would actually sign.
  7. 7Day 0 and day 2 — build a group from nothing, then upgrade under live traffic.

Licensing

Free for most people. One-time for everyone else.

Bloodraven is source-available under the Business Source License. Prices, the eligibility table and the checkout links are on one page, so there is exactly one place to read what you owe.

  • Free forever for individuals, non-commercial use, non-profits and companies under $1M annual revenue
  • Free forever for dev, test, staging, CI and evaluation, at any scale
  • Production at a company over $1M annual revenue is a one-time licence, not a subscription
  • No activation, no licence server, no feature gating — the software is fully functional without a key
  • Every version converts to Apache 2.0 two years after it is published

Bloodraven · ShipStream

Try it in two minutes. Then break it on purpose.

Install the operator, create a failover group, then kill a site on purpose and watch it recover without you. Everything on this page runs on a laptop first.