Kubernetes operator for MySQL
Give your MySQL topology a survival instinct.
Bloodraven runs MySQL async replication failover groups across sites. It owns detection, fencing, promotion, DNS steering, clone bootstrap and backups — so losing an entire site is a status change, not an incident.
operator ready · watching 3 sites · press Kill the active site to watch a failover.
Scripted replay of the documented promotion sequence. Kill the site and read the clock.
- ~6s
- to declare a site unreachable (3 failed 2s polls)
- RPO 0
- on a planned switchover, by construction
- 47+
- scripted chaos scenarios run in CI
- 1
- CRD, one controller, one reconcile loop
Three ways in
Sixty pages of docs, one prompt
Don't read 60 pages. Just ask.
The assistant is grounded in this entire documentation set — the CRD reference, every runbook, the failure-mode matrix and the log schema. Ask it in your own words and it answers with links to the exact page.
Start here
Day-2+ Diagnostic Copilot
Meet Bloodraven Doctor. AI diagnostics in the CLI.
Give your AI coding assistant deep diagnostic vision over live MySQLFailoverGroup clusters. bloodraven-doctor performs 60-second non-destructive triage, detects GTID divergence, audits keyring status, and crafts safe step-by-step remediation plans.
- 60-Second Fast Triage
- Non-destructive assessment of quorum, active primaries, split-brain status, and replication lag.
- GTID Divergence Auditor
- Compares executed transaction sets across sites to detect errant writes and guide safe recloning.
- Keyring Escrow Safety
- Enforces strict safety guards to protect master encryption keys from loss during pod restarts.
- Runbook-Grounded Plans
- Produces step-by-step remediation commands with clear blast radius and RPO trade-offs.
1. Monitor catchup: kubectl bloodraven status orders -n orders2. Preflight failover: kubectl bloodraven promote orders --to=pdx (blocked until lag=0s)Detection · fencing · promotion
Losing a site should be boring.
The operator watches every site on a two-second poll, decides with a small documented state machine, and executes the same promotion sequence every time.
Automated DNS failover
Be genuinely geo-redundant without a global load balancer in front of your database. Bloodraven writes an external-dns DNSEndpoint with your hostname and TTL, and moves it to the promoted site as part of the promotion sequence — no cross-region LB bill, no anycast VIP, no proxy hop on every query.
external-dns DNSEndpoint · your hostname · your TTLRead how it works ↗02SAFESplit-brain safe, and tested that way
Two sites never both accept writes. The operator fences the old primary, and each sidecar self-fences with super_read_only when it can reach neither the operator nor its peer. GTID divergence is detected, reported in divergentGtid, and blocks an unsafe rejoin until a human decides.
operator fencing · sidecar super_read_only · divergentGtidRead how it works ↗03ROLLZero-downtime updates
OrderedUpdate upgrades the standby first, fails over to it, then upgrades the old active — the direction MySQL's rolling-upgrade contract requires. Node taints and the placement contract keep application workloads on the same site as the writable MySQL as the primary moves.
OrderedUpdate · standby first · placement contractRead how it works ↗Bootstrap, scale out, encrypt, provision, back up
Data that looks after itself.
Bootstrap, read scale-out, encryption, tenant provisioning and backup are part of the operator, not five more systems you have to wire together.
Clone-based bootstrap
New replicas seed themselves with MySQL's clone plugin and pick up replication with GTID auto-positioning. No mysqldump window, no snapshot juggling, no manual data transfer. The same path repairs a site whose PVC was lost, or one you deliberately reclone after divergence.
CLONE INSTANCE FROM 'replicator'@'orders-iad:3306' …
CHANGE REPLICATION SOURCE TO SOURCE_AUTO_POSITION = 1;
START REPLICA;Explore →AESData-at-rest encryption, for free
InnoDB tablespace encryption on ordinary PVCs, using the GPL keyring component that ships with MySQL Community Edition. No Oracle Enterprise licence, no encrypted CSI storage class, and the master key never lands on the data PVC or a worker-node disk. Rotation, sealing and escrow are handled by the operator.
spec:
encryptionAtRest:
enabled: trueExplore →RORead-only replicas
Append a read-only site and it follows whichever site is active — never promoted, never a planned-failover target, never a DNS target, never a clone donor. It gets its own client Service and its own mysqlConf, so you can size it for reporting, analytics or a CDC tap without touching the write path. Fall behind readOnlyMaxLagSeconds and its endpoint sheds until it catches up; lose its data and it reclones itself.
- name: reader
role: read-only
zone: us-east-1bExplore →CRDTenant databases as resources
Declare a schema, its owner and any extra principals — a per-tenant SELECT-only support reader, say — as a MysqlDatabase, and the operator runs the SQL on whichever site is primary. Every account is Secret-backed, scoped to its own schema and, if you like, to a list of source IPs; rotation is a Secret write, and removing an entry revokes and drops it. Nothing but Bloodraven holds the admin credential.
users:
- secretName: support-ro-mysql
privileges: [SELECT]
hosts: [35.1.2.3, 35.4.5.6]
resourceLimits:
maxUserConnections: 5Explore →Backup and restore. The whole nine yards.
Not a cron job that shells out to mysqldump. Backup is a first-class part of the operator, from the schedule all the way through to proving the artifact can actually be restored.
- S3 or PVC — object storage for durability, PVC for labs
- Scheduled and on-demand — a MysqlBackup whenever you need one
- Structured retention — plus exponential-backoff retries on failure
- Point-in-time recovery — binlog archiving between full dumps
- Encrypted artifacts — application-level, independent of the store
- Prometheus metrics — every run, every failure, every sweep
- Automatic cleanup — deleting the object removes its artifacts
- Verification — load the dump into a throwaway MySQL and prove it
- initFromBackup — bootstrap a brand-new failover group from a dump
- 01Dumpconsistent, scheduled or on demand
- 02RetainS3 or PVC · structured retention
- 03Verifyrestore into a throwaway MySQL
- 04BootstrapinitFromBackup into a new group
No quorum. No coordinator.
One controller. One CRD. One loop.
There is no distributed consensus in Bloodraven, no coordinator to keep quorum, and no second system to reconcile against. A single reconcile loop reads the observed topology and writes the decision — which is why the failure modes fit on one page.
- No coordination problem
- Nothing to elect, nothing to split. State lives in the CR's status and in MySQL itself.
- The data plane doesn't need the operator
- A healthy primary and replica keep serving reads and writes with zero operator involvement — it is on the detection path, not the request path.
- Correct even while the operator is down
- Sidecars self-fence when they can reach neither the operator nor their peer, so no split brain is possible while the control plane is missing.
- Every failure mode is written down
- A documented matrix of faults, what the operator does, how long it takes, and what it costs you.
1apiVersion: shipstream.io/v1alpha1 2kind: MysqlFailoverGroup 3metadata: 4 name: orders 5spec: 6 dns: 7 hostname: orders.az.example.com 8 ttl: 60 9 replication:10 readOnlyMaxLagSeconds: 3011 sites:12 - name: iad13 zone: us-east-1a14 lbIP: 10.0.1.115 - name: pdx16 zone: us-west-2a17 lbIP: 10.0.2.118 - name: reader19 role: read-only20 zone: us-east-1bThe app finds out in milliseconds
Your application finds out immediately.
A failover the app never notices is the point. Bloodraven pushes topology changes to connected clients and moves your cache along with the database.
Push the topology the instant it changes.
A WebSocket stream publishes the full topology of every failover group the moment it changes, so an app can force a pool reconnect on promotion instead of discovering it through a wall of write errors. A REST status API and Prometheus metrics expose the same state.
const ws = new WebSocket('ws://bloodraven:8082/ws/status')
ws.onmessage = e => pool.reconnectIfPrimaryMoved(JSON.parse(e.data))Integrate your app →Cache and sessions move with the primary.
Turn on spec.dragonfly and Bloodraven runs a Redis-compatible cache and session store per site that follows MySQL. Cache and session continuity, not another failover to operate.
- 1
WAITreplica offset catches up - 2
REPLTAKEOVERpromote without a restart - 3
CLIENT KILLold-master clients reconnect to the active endpoint
Deployment API · /deploy/v1
Migrations that know the ground moved.
A deploy pipeline asks Bloodraven for a lease before it touches the schema. The lease serializes migrations across the group, holds planned failovers still while it is renewed, and is revoked the instant the topology generation changes — so a migration can never keep running on a primary it did not start on.
Two deploys collide. One migration per group. The second deploy gets 409 held, with the holder, and waits its turn.
deploy client ready · two deploys collide · press Play to watch the API arbitrate.
MUTEXOne migration per group- A second deploy gets 409 held with the holder's operation ID and expiry. Two pipelines cannot race the same schema, and a crashed one stops blocking when its TTL lapses.
HOLDPlanned disruption waits- A renewed failover-hold defers planned promotion, ordered updates and restore-in-place with phase Deferred and reason DeploymentHold. It never delays emergency failover or primary fencing.
FENCERevoked, not reconnected- Every lease is stamped with topologyGeneration. Any authoritative site change bumps it and the next renewal returns 409 revoked, so the client stops instead of continuing DDL against the new primary.
RBACNo cluster-admin for CI- Callers authenticate with a projected ServiceAccount token and are authorized from MysqlDatabase.spec.deploymentClients. No permission to patch the group, no Secrets, no MySQL credentials, TLS only.
Run it on your laptop first
Proof, not promises.
Every claim on this page is exercised against a real Kubernetes cluster — nightly in CI, and on your laptop in about two minutes.
Primary kills. Partitions. Self-fencing. GTID divergence. Data wipes.
47+ scripted chaos scenarios, including operator crashes mid-failover, rolling updates, Dragonfly failover, and backup and PITR verification. Each one states a hypothesis, injects a real fault, asserts on operator behaviour and captures full forensics when it fails. A smoke subset gates every release before artifacts are published.
make chaos-run SCENARIO=06-self-fence-isolated-primary
make chaos-run-all-profile PROFILE=smokeA whole cluster. One script.
Two MySQL sites, a live dashboard, a counter app that writes through the failover hostname, DNS visualisation and a chaos menu. Break it on purpose and watch the whole promotion happen in front of you.
$./playground/setup.sh▍
Train the on-call rotation
Bloodraven in Production.
Train your on-call team on a real cluster, not on slides. Hold a site down and time the promotion. Audit the exact transactions an emergency failover cost you. Wire an application that survives one. Roll a MySQL upgrade underneath live traffic.
- 7
- units
- 28
- topics
- 35
- quizzes and tests
Every number in the course traces to a source: operator code, the shipped CRDs, recorded chaos-run forensics, and the MySQL manual.
- 1Meet the group — stand up a real three-site failover group and read its status.
- 2How the operator decides — predict its next move from a status dump alone.
- 3Emergency failover end to end — time the promotion, audit what it cost.
- 4Where failover meets your application — pools, reconnects and stale reads.
- 5When the world misbehaves — self-fencing, split brain, five kinds of partition.
- 6Backups, disaster recovery and a go-live checklist you would actually sign.
- 7Day 0 and day 2 — build a group from nothing, then upgrade under live traffic.
Licensing
Free for most people. One-time for everyone else.
Bloodraven is source-available under the Business Source License. Prices, the eligibility table and the checkout links are on one page, so there is exactly one place to read what you owe.
- Free forever for individuals, non-commercial use, non-profits and companies under $1M annual revenue
- Free forever for dev, test, staging, CI and evaluation, at any scale
- Production at a company over $1M annual revenue is a one-time licence, not a subscription
- No activation, no licence server, no feature gating — the software is fully functional without a key
- Every version converts to Apache 2.0 two years after it is published
Bloodraven · ShipStream
Try it in two minutes. Then break it on purpose.
Install the operator, create a failover group, then kill a site on purpose and watch it recover without you. Everything on this page runs on a laptop first.

