Fencing forensics

Goal

(Optional — see the brief for the shorter jq route and what skipping it costs you.) Build a timeline tool that reads a bundle of operator and sidecar logs from an injected fault and reports every fence, which rule caused it, and whether it was correct — so that a read-only site you did not expect becomes a question you can answer from evidence rather than a guess.

Unit 5 — When the world misbehaves · project · code-notebook · Python

Optional. The heaviest project in the course, and the most specialised. Nothing later depends on it. Skip it if you are short of time, and take the deliverable instead of the code: a read-only site is explained by exactly one of three causes — rule #1 (a live operator disagrees about the active site), rule #2 (operator and every peer silent past leaseTimeout), or the startup safety net (the pod never got permission) — and the log prefixes tell them apart. safety net: means never allowed to write; SELF-FENCING: / SELF-FENCED: means was writing and lost the argument. Being able to say which, from a log bundle, is the whole objective.

Goal

Build a timeline tool that reads a bundle of operator and sidecar logs from an injected fault and reports every fence, which rule caused it, and whether it was correct — so that a read-only site you did not expect becomes a question you can answer from evidence rather than a guess.

You will:

  1. Attribute every fence in a log bundle to the rule or safety net that caused it.
  2. Judge from evidence whether a fence was correct, premature, or never happened at all.

Everything runs against the JSON-lines fixtures in tests/fixtures/. You do not need a live cluster, and neither does the grader.


How this works

A site on playground is read-only and you did not put it there. That is the whole problem. Three different mechanisms could have done it — the sidecar’s rule #1, the sidecar’s rule #2, or the startup safety net — plus three ways the operator fences a site itself, and they mean completely different things. Rule #1 is a pod that was told somebody else is active. Rule #2 is a pod that could not reach anyone. The safety net is a pod that was never allowed to write in the first place. Same read-only outcome, three different incidents.

You have the evidence. Both binaries write structured JSON to stdout, and site/content/docs/8.observability/7.log-schema.md is a stability contract: filter on msg and those strings will not move without a deprecation note. So the forensics are mechanical, once you know the vocabulary.

You are building brfence. It takes one bundle directory:

tests/fixtures/partition-a/
  operator.jsonl        the operator's operational (slog) stream
  sidecar-iad.jsonl     one file per site's bloodraven-sidecar container
  sidecar-pdx.jsonl
  sidecar-reader.jsonl

and prints one timeline of every fence, with its cause and a verdict, plus the writable sites nobody fenced at all. Exit 0 when everything checks out, 1 when it does not.

Three bundles ship with the project, all from playground:

BundleWhat it is
tests/fixtures/partition-aA shape-A partition. iad is isolated, pdx is promoted in 12.0 s, iad self-fences on the lease, the reader pod restarts, and iad comes back writable for two seconds before its own monitor catches it.
tests/fixtures/split-brain-tier3A split brain on a group with no sitePriorities — tier 3, alert only. One sidecar log was never collected. One fence is not supported by its own evidence.
tests/fixtures/decoysThirteen records, eleven of which contain the string fenc, and exactly one fence.

Your tasks

TODO A — classify(rec, stream). Return one of six cause ids for a record that is a fence decision, and None for everything else. The vocabulary is in the fence-vocabulary reference. Match the whole msg. SELF-FENCING: is a stable prefix, not a synonym for “a fence happened” — it also heads SELF-FENCING: killed app connections (a follow-up) and SELF-FENCING FAILED: could not set super_read_only (a fence that did not land).

TODO B — fenced_site(rec, cause, file_site). Return the site that got fenced. Records disagree about how to say it: rule-1 and the operator’s old-primary and non-promotable lines carry site, the split-brain line carries fencedSite, and the rule-2 and two of the three safety-net lines carry no site at all. The loader hands you file_siteiad for sidecar-iad.jsonl, None for the operator stream — for exactly that case.

TODO C — judge(rec, cause, site). Return "correct" or "premature" using the verdict-table reference. Rule-2 is the one that earns its keep: the lease fence may only fire when the operator and every peer have been silent for the whole leaseTimeout, so subtract bloodravenLastOk and latestPeerOk from the record’s own time and compare. A peer that answered four seconds ago means the fence should not have happened. parse_time() and parse_duration() are already written.

TODO D — unfenced_writable_sites(records, fences). The operator emits msg="ALERT" with a message field, and a split-brain alert reads exactly SPLIT BRAIN: 2 sites are writable (iad, pdx). Take the names from inside the parentheses of every such alert. Drop any site that has a fence event anywhere in the bundle, and drop the site named by the most recent failover complete promotedSite at or before the alert — that site holds primary authority and fencing it would be the bug. Return (site, alert_message) pairs, each site once, in first-seen order.

What the scaffolding is for

Everything that is not forensics is already wired, so you spend your time on the discriminations rather than on plumbing:

Expected output

FENCE TIMELINE — playground (bundle: split-brain-tier3)
  11 records scanned, 3 fence events

  2026-08-12T14:01:03.505Z  reader   safety-net     correct    error=Get "http://…/active-site": context deadline exceeded
  2026-08-12T14:02:15.130Z  reader   non-promotable correct    site=reader role=read-only
  2026-08-12T14:04:20.880Z  iad      rule-2         premature  bloodravenLastOk=… latestPeerOk=… leaseTimeout=20s

UNFENCED WRITABLE SITES
  pdx  — writable per ALERT "SPLIT BRAIN: 2 sites are writable (iad, pdx)", no fence event in this bundle

VERDICT: 3 fences, 1 premature, 1 unfenced writable site

tests/fixtures/partition-a gives three fences, all correct, (none) unfenced, exit 0. tests/fixtures/decoys gives exactly one fence, exit 0.

Rules


Reference: The fence vocabulary

Six causes. Eight msg strings. These are the stable identifiers from the log-schema contract — match the whole string, in the stream named.

Cause idStreammsg (verbatim)
rule-1sidecarSELF-FENCING: topology mismatch — operator-authoritative active site disagrees with our site, setting super_read_only=ON
rule-2sidecarSELF-FENCING: Bloodraven and every peer unreachable beyond lease timeout, setting super_read_only=ON
safety-netsidecarsafety net: could not query active site, staying fenced
safety-netsidecarsafety net: no active site reported by operator, staying fenced
safety-netsidecarsafety net: confirmed standby site, staying fenced
split-brainoperatorsplit-brain auto-resolve: fencing non-preferred site per spec.splitBrainPolicy.sitePriorities
old-primaryoperatorfencing returning old primary (split brain after failover)
non-promotableoperatorfenced writable non-promotable site

The em dash in the rule-1 string is U+2014, not a hyphen.

Not a fence decision

Every one of these contains fenc and none of them is an event you count. They are in the fixtures on purpose.

msgWhat it actually is
SELF-FENCED: super_read_only=ON has been set, only Bloodraven can restoreThe status line that follows a fence. Counting it doubles every self-fence.
SELF-FENCING FAILED: could not set super_read_onlyA fence that did not land. The sidecar retries next tick.
SELF-FENCING: killed app connectionsEviction after a fence that already succeeded.
SELF-FENCING: failed to kill connections after fencingThe fence holds; eviction was incomplete.
SELF-FENCING: super_read_only write failed but the fence is in place; skipping connection evictionThe write errored but a follow-up read shows it landed.
failed to fence old primary (may be unreachable)Step 1 of the failover sequence missing an unreachable host. It only warns.
failed to fence returning old primaryThe operator’s fence write was rejected.
fencing: adopted active-site view from peerA peer relayed a fresher topology view. This is what drives rule #1, not rule #1.
fencing: MySQL is writable after prior self-fence; rearming monitorThe opposite of a fence.
fencing: could not check read_only status / fencing: could not confirm whether the super_read_only write landedProbe failures.
re-asserting fenced promoted primary: …The operator removing a fence.

The fields you need

Common to every operational record: time (RFC3339Nano, always Z), level, msg, fg. Sidecar records also carry pod. Then, per cause:

CauseSite named byEvidence fields
rule-1siteauthoritativeActiveSite, observedAt
rule-2nothing — use the filenamebloodravenLastOk, latestPeerOk, peers, leaseTimeout
safety-netsite, but only on confirmed standby siteactiveSite, error
split-brainfencedSitewinner
old-primarysite
non-promotablesiterole

Reference: The verdict table

A fence is correct when the record’s own fields support the rule that fired, and premature when they contradict it. You are not second-guessing the operator’s judgement — you are checking whether the evidence it logged actually meets the condition it claims to have met.

Causecorrect whenWhy
rule-1authoritativeActiveSite is non-empty and is not the fenced siteRule #1 is exactly “the authoritative active site is somebody else”. If the record names this same site, the fence has no premise.
rule-2time − bloodravenLastOk ≥ leaseTimeout and (latestPeerOk absent or time − latestPeerOk ≥ leaseTimeout)The lease fence requires the operator and every peer silent for the full window. One reachable peer keeps the primary writable. A missing latestPeerOk means no peer ever answered — that is silence, so it counts as met.
safety-netalwaysThe safety net fails closed by staying fenced. It never takes writability away from a site that had it, so it cannot fire too early.
split-brainwinner is non-empty and is not the fenced siteTier 2 fences the losers and re-promotes the winner. A record that fences its own winner is broken.
old-primaryalwaysThe operator fences every writable site except the one holding live primary authority. The record carries nothing that could contradict it; read the timeline around it.
non-promotablerole is not primary-candidatePromotability is exactly role == primary-candidate. A writable dr-only or read-only site is never authoritative and is fenced on every poll. Anything else in role means the wrong site was fenced.

leaseTimeout arrives as slog’s duration rendering — the string "20s". parse_duration() handles it. The shipped defaults are leaseTimeout: 20s and peerCheckInterval: 5s, so a lease covers four consecutive ticks of silence.

The fence that never happened

There is a fourth possibility the timeline alone will not show you: nobody fenced anything. On a group with no splitBrainPolicy.sitePriorities and no usable failover history, tier 3 is alert only — the operator emits SPLIT BRAIN: n sites are writable (…) every poll and takes no action, by design and by field documentation. That is a correct operator and an unresolved split brain at the same time, and the only way to see it in a bundle is to notice which writable sites never appear in the fence timeline.

Two sites you must not report as unfenced: one that was fenced later in the bundle, and the site named by the most recent failover complete promotedSite, which is the one legitimately holding primary authority.


Steps


Grading

Graded three ways: the steps above, the human rubric in rubric.md, and the four machine test cases in project.json — mirrored in tests/test_brfence.py, which you can run yourself at any time.

Steps

How it is graded

CriterionWhat earns itWeight
Fence vocabulary and classificationclassify() maps whole msg strings to the six cause ids rule-1, rule-2, safety-net, split-brain, old-primary, non-promotable, and returns None for everything else. Full marks require exact-string matching: any solution that prefix-matches SELF-FENCING: or substring-matches fenc scores at most half, because it counts SELF-FENCING: killed app connections, SELF-FENCING FAILED: and SELF-FENCED: as fences. Deduct if a sidecar-only cause is accepted from the operator stream or vice versa.30
Site attribution across inconsistent record shapesfenced_site() returns the right site for all six causes, not just the ones carrying a site key. Full marks require handling the rule-2 and safety-net records that name no site (fall back to the filename-derived file_site) and the operator's fencedSite key. A solution that emits ?, None, or the pod name for any fence in the three fixtures loses most of this criterion.25
Verdicts and the missing fence, judged from the recordjudge() implements the verdict table and, critically, computes the rule-2 verdict from time - latestPeerOk and time - bloodravenLastOk against leaseTimeout rather than assuming a lease fence is self-justifying. unfenced_writable_sites() parses the SPLIT BRAIN: alert, subtracts sites that were fenced and the site holding primary authority, and reports each remaining site once even when the alert repeats every poll. Deduct for a hardcoded verdict, a hardcoded site list, or reporting the promoted site as unfenced.25
Craft — clarity, structure, and behaviour on imperfect bundlesThe four functions stay separate and single-purpose; no fixture path, site name, or expected count is hardcoded. The tool survives a bundle with a missing sidecar file, a record with no site key, an unknown msg, and an unexpected extra field without crashing, and the exit code is 1 exactly when there is a premature fence or an unfenced writable site. Comments explain the discriminations that are easy to get wrong — why SELF-FENCED: is not counted, why a lease fence needs checking — rather than restating the code.20
Total100

Test cases

Test Checks Expected Weight
canonical_timeline Correctness on the canonical input. Runs the tool against tests/fixtures/partition-a, a shape-A partition on playground: iad self-fences on rule-2 while isolated, the restarted reader sidecar stays fenced by the startup safety net, and iad self-fences again on rule-1 after its mysqld comes back writable. Asserts three fence events with the right site, cause and verdict in time order, no premature fences, no unfenced writable sites, and exit 0. PASS 40
awkward_bundle_generality Generality on an awkward bundle. tests/fixtures/split-brain-tier3 has no sidecar-pdx.jsonl at all, two fence records that carry no site key, a repeated SPLIT BRAIN alert, an operator fence of a read-only reader, and a rule-2 fence whose own latestPeerOk sits four seconds inside the twenty-second lease window. Asserts the rule-2 line is premature, that pdx and only pdx is reported unfenced, and that the tool exits 1. PASS 25
decoy_lines_are_not_fences Catches a shortcut: matching msg by substring or by the SELF-FENCING: prefix instead of by the whole string. tests/fixtures/decoys holds thirteen records, eleven containing fencSELF-FENCED:, SELF-FENCING FAILED:, SELF-FENCING: killed app connections, failed to fence old primary, fencing: adopted active-site view from peer — and exactly one real fence decision. A substring matcher reports eleven, a prefix matcher four; the test demands exactly one, iad/rule-1/correct, and exit 0. PASS 20
stable_msg_vocabulary_in_source Structural. Asserts the required construct is present in brfence.py: the two SELF-FENCING: msg strings and all three safety net: … staying fenced strings appear verbatim (whitespace and implicit string concatenation normalised), and the source reads bloodravenLastOk, latestPeerOk and leaseTimeout — proof the rule-2 verdict was computed from the record rather than assumed. PASS 15

Starter files

1 files

brfence.py

                    #!/usr/bin/env python3
"""brfence — fencing forensics for one MysqlFailoverGroup log bundle.

Reads a directory of JSON-lines logs collected from a single injected fault on
`playground` and reports every fence, what caused it, and whether the record's own
evidence supports it.

    python3 brfence.py tests/fixtures/partition-a

Bundle layout:

    operator.jsonl          the operator's operational (slog) stream
    sidecar-<site>.jsonl    one file per site's bloodraven-sidecar container

Everything except the four TODOs is already wired: bundle loading, timestamp
and duration parsing, ordering, the report format, and the exit code.

Do not change the report format. The tests read it.
"""

from __future__ import annotations

import argparse
import json
import sys
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from pathlib import Path

# Fields that appear on nearly every record and carry no forensic weight.
# The report prints everything else as evidence.
BORING_FIELDS = ("time", "level", "msg", "fg", "pod")


# --------------------------------------------------------------------------
# TODO A — the fence vocabulary.
#
# Return the cause id for a record that is a fence *decision*, or None for
# everything else. The six cause ids and the exact `msg` strings that map to
# them are in the brief. Key on the whole `msg` string: the `SELF-FENCING:`
# prefix is stable but it is not a synonym for "a fence happened".
#
#   rec     the decoded JSON record
#   stream  "operator" or "sidecar" (which file it came from)
# --------------------------------------------------------------------------
def classify(rec: dict, stream: str) -> str | None:
    return None  # TODO A


# --------------------------------------------------------------------------
# TODO B — which site got fenced.
#
# Not every fence record carries a `site` key, and the operator's split-brain
# line calls it something else. Fall back to `file_site` (the site name taken
# from the sidecar filename) when the record does not name a site itself.
# Return a site name, never None.
#
#   file_site  "iad" for sidecar-iad.jsonl, None for operator.jsonl
# --------------------------------------------------------------------------
def fenced_site(rec: dict, cause: str, file_site: str | None) -> str:
    return "?"  # TODO B


# --------------------------------------------------------------------------
# TODO C — was the fence supported by its own evidence?
#
# Return "correct" or "premature". The verdict table is in the brief. The
# interesting one is rule-2: it may only fire when the operator *and* every
# peer have been silent for the whole `leaseTimeout`, so compare the record's
# `bloodravenLastOk` and `latestPeerOk` against its `time`.
#
# parse_time() and parse_duration() below are already written for you.
# --------------------------------------------------------------------------
def judge(rec: dict, cause: str, site: str) -> str:
    return "correct"  # TODO C


# --------------------------------------------------------------------------
# TODO D — writable sites that nobody fenced.
#
# The operator emits `msg="ALERT"` with a `message` field. A split-brain alert
# reads exactly:  SPLIT BRAIN: 2 sites are writable (iad, pdx)
#
# Take the site names from inside the parentheses of every such alert, then
# drop:
#   * any site that has a fence event anywhere in this bundle, and
#   * the site named by the most recent `failover complete` `promotedSite` at
#     or before the alert — that site holds primary authority; fencing it
#     would be the bug.
#
# Return a list of (site, alert_message) pairs, each site at most once, in the
# order the sites first appear.
#
#   records  list of (rec, stream, file_site), already sorted by time
#   fences   list of Fence, already built from TODOs A-C
# --------------------------------------------------------------------------
def unfenced_writable_sites(records: list, fences: list) -> list:
    return []  # TODO D


# ==========================================================================
# Everything below is wired. You should not need to change it.
# ==========================================================================


@dataclass
class Fence:
    when: datetime
    raw_time: str
    site: str
    cause: str
    verdict: str
    rec: dict = field(repr=False)


def parse_time(value: str) -> datetime:
    """RFC3339Nano as the binaries emit it, e.g. 2026-08-12T09:14:08.113Z."""
    text = value.strip()
    if text.endswith("Z"):
        text = text[:-1] + "+00:00"
    if "." in text:
        head, _, tail = text.partition(".")
        digits = "".join(c for c in tail if c.isdigit())
        offset = tail[len(digits):]
        text = f"{head}.{digits[:6]:<06}{offset}"
    return datetime.fromisoformat(text).astimezone(timezone.utc)


def parse_duration(value: str) -> timedelta:
    """slog's default time.Duration rendering: "20s", "500ms", "5m"."""
    text = str(value).strip()
    for suffix, factor in (("ms", 0.001), ("s", 1.0), ("m", 60.0), ("h", 3600.0)):
        if text.endswith(suffix):
            return timedelta(seconds=float(text[: -len(suffix)]) * factor)
    return timedelta(seconds=float(text))


def load_bundle(bundle: Path) -> list:
    """Read every *.jsonl in the bundle, sorted by event time.

    Returns a list of (record, stream, file_site) tuples. `stream` is
    "operator" or "sidecar"; `file_site` is the site name taken from a
    sidecar-<site>.jsonl filename, or None for the operator stream.
    """
    rows = []
    for path in sorted(bundle.glob("*.jsonl")):
        name = path.stem
        if name == "operator":
            stream, file_site = "operator", None
        elif name.startswith("sidecar-"):
            stream, file_site = "sidecar", name[len("sidecar-"):]
        else:
            continue
        for lineno, line in enumerate(path.read_text(encoding="utf-8").splitlines()):
            line = line.strip()
            if not line:
                continue
            try:
                rec = json.loads(line)
            except json.JSONDecodeError as exc:
                print(f"{path.name}:{lineno + 1}: skipping unparseable line: {exc}",
                      file=sys.stderr)
                continue
            if "time" not in rec or "msg" not in rec:
                continue  # controller-runtime (zap) noise, not the operational stream
            rows.append((rec, stream, file_site, path.name, lineno))
    rows.sort(key=lambda r: (parse_time(r[0]["time"]), r[3], r[4]))
    return [(rec, stream, file_site) for rec, stream, file_site, _, _ in rows]


def group_name(records: list) -> str:
    for rec, _, _ in records:
        fg = rec.get("fg")
        if fg:
            return fg.split("/")[-1]
    return "playground"


def evidence(rec: dict) -> str:
    parts = []
    for key, value in rec.items():
        if key in BORING_FIELDS:
            continue
        parts.append(f"{key}={value if isinstance(value, str) else json.dumps(value)}")
    return " ".join(parts)


def plural(n: int, word: str) -> str:
    return f"{n} {word}" if n == 1 else f"{n} {word}s"


def report(group: str, bundle: Path, records: list, fences: list, unfenced: list) -> int:
    print(f"FENCE TIMELINE — {group} (bundle: {bundle.name})")
    print(f"  {plural(len(records), 'record')} scanned, {plural(len(fences), 'fence event')}")
    print()
    if not fences:
        print("  (no fence events in this bundle)")
    for f in fences:
        print(f"  {f.raw_time:<26}{f.site:<9}{f.cause:<15}{f.verdict:<11}{evidence(f.rec)}")
    print()
    print("UNFENCED WRITABLE SITES")
    if not unfenced:
        print("  (none)")
    for site, alert in unfenced:
        print(f"  {site}  — writable per ALERT \"{alert}\", no fence event in this bundle")
    print()
    premature = [f for f in fences if f.verdict != "correct"]
    print(f"VERDICT: {plural(len(fences), 'fence')}, {len(premature)} premature, "
          f"{plural(len(unfenced), 'unfenced writable site')}")
    return 1 if (premature or unfenced) else 0


def main(argv: list | None = None) -> int:
    ap = argparse.ArgumentParser(description="Fencing forensics for a Bloodraven log bundle.")
    ap.add_argument("bundle", type=Path, help="directory holding operator.jsonl and sidecar-*.jsonl")
    args = ap.parse_args(argv)

    if not args.bundle.is_dir():
        print(f"brfence: {args.bundle} is not a directory", file=sys.stderr)
        return 2

    records = load_bundle(args.bundle)
    if not records:
        print(f"brfence: no operational log records found in {args.bundle}", file=sys.stderr)
        return 2

    fences = []
    for rec, stream, file_site in records:
        cause = classify(rec, stream)
        if cause is None:
            continue
        site = fenced_site(rec, cause, file_site)
        fences.append(Fence(
            when=parse_time(rec["time"]),
            raw_time=rec["time"],
            site=site,
            cause=cause,
            verdict=judge(rec, cause, site),
            rec=rec,
        ))

    unfenced = unfenced_writable_sites(records, fences)
    return report(group_name(records), args.bundle, records, fences, unfenced)


if __name__ == "__main__":
    sys.exit(main())

                  

Projects are not auto-graded here. The rubric and the test cases above are the grading contract — run them yourself on your own machine.

Erase saved progress?

This erases all quiz scores, reading progress, project checklists, and your name on the certificate. It cannot be undone, and it affects only this course in this browser.