49 KiB
Session 0007.0 — Transcript
Date: 2026-05-27 Goal: Execute Slice 4 of the
ohm-rfc-app-flotillaSPEC §20 slicing plan — full deploy gesture +/api/healthverification. Per §20.4: add thedeploysandhealth_snapshotstables; ship the real SSH gesture viagcloud compute sshper §8.1 with fail-stop semantics per §8.2; the §9 probe loop with the §9.2 polling shape and the §9.3 verdict table; the §8.3 concurrency lock; the §8.1 step 6 atomic.envwrite on the VM; new CLI verbsdeploy <name>,deploy abort,deploy status,deploy watch,deploy log. Exit criterion: the live gesture recordsoutcome=succeededand/api/healthreports the deployed version, ships as v0.4.0.Outcome: Slice 4 code shipped at v0.4.0 on the working tree — not yet committed, not yet tagged, and the live first-deploy gesture against the OHM VM (§20.4 final exit criterion) is the operator gesture this release is gated on. All 132 tests pass locally (74 baseline → 132, +58 for Slice 4). New surfaces: migration
004_deploys_health_snapshots.sql(with a SQLite partial-unique- index implementing the §8.3 in-flight lock);ohm_rfc_app_flotilla/ ssh.py(a thin DI-friendlygcloud compute sshwrapper);ohm_rfc_app_flotilla/health.py(one-shot probe + §9 probe loop with the §9.3 verdict table verbatim);ohm_rfc_app_flotilla/ deploy_log.py(CRUD fordeploysandhealth_snapshots, phase appender, abort path,redact()for §10.2);ohm_rfc_app_flotilla/ deploy.py(the 9-phase orchestration with secret-byte scrubbing via a_ResolvedSecretsbytearray context). CLI surface extended to the full §12.1 verb inventory.pin checknow compares against the most recent successful deploy.Five findings worth flagging permanently:
- Click
invoke_without_command=Truedoes not compose with a positional[DEPLOYMENT]argument on the group. Slice 4 neededdeploy <name>(a verb) anddeploy abort <name>(a subcommand) to coexist per SPEC §12.1. The naive pattern (@main.group("deploy", invoke_without_command=True)+@click.argument("deployment", required=False)) parsesdeploy ohm-rfc-app --dry-runas DEPLOYMENT=ohm-rfc-app, then tries to invoke a subcommand named--dry-runand fails. Resolution: a custom_DeployGroup(click.Group)subclass overridesresolve_commandto route any non-subcommand first arg to a hiddenrunsubcommand. The §12.1 surface is preserved without compromise. Implementation note now sits incli.pyso a future maintainer doesn't re-discover it.- SQLite partial unique indexes are the right tool for the §8.3 in-flight lock. A
CREATE UNIQUE INDEX … WHERE outcome = 'in_progress'overdeploys(deployment_id)makes the database refuse a second in-flight row directly;deploy_log.open_deploycatchessqlite3.IntegrityErrorand translates it intoConcurrencyErrorwith the explicit "deploy already in flight" message. No advisory locking, no application-level mutex; the schema enforces the invariant.- The
/api/health§9.3 verdict table maps cleanly to aVerdictenum + a flatevaluate(reading, target_version)function. Each row of the table becomes one branch; the probe loop reads the return value and decides whether to exit immediately (SUCCESS, ENDPOINT_MISSING), keep polling within budget (MISMATCH, DEGRADED), or back off and retry (TRANSIENT). Connection-failure readings (http_status is None) fold into TRANSIENT — matches §9.3 row 4 "connection refused → transient".- Secret-byte scrubbing is best done with a context-managed
bytearraycontainer._ResolvedSecretsindeploy.pyholds each resolved payload as abytearray;__exit__(always called, including during exception unwind) overwrites every byte with0x00and clears the dict. The env body itself is also abytearrayzeroed in afinallyafter the VM-side write completes. Python's immutablestrmakes truly-secure scrubbing impossible, but the bytearray path stays in our control end-to- end and the GC never sees the payload as a managed object.- Atomic
.envwrite via single SSH command piping body tocat > tmp && mv -f tmp final. Both files live in the same directory, so the rename isrename(2)— atomic on the local filesystem.umask 077inside thesh -cbody makes the temp file mode 0600;sudo -u <user>makes it owned by the service user. The stdin pipe means the env body never lands on the operator's disk; one bytearray, one network hop, scrubbed at scope exit.No SPEC corrections this slice. The implementation discoveries above are Click and SQLite mechanics, not contradictions of the spec — §8.3 reads cleanly against the partial-unique-index; §9.3 reads cleanly against the
Verdictenum; §8.1 step 6 reads cleanly against the bytearray + atomic-rename pattern. Per §19.3 rule 2, no edits required.
Turn 1 — User
Let's do slice 4
Turn 1 — Assistant
Loaded the project memory and re-read the relevant SPEC sections.
MEMORY.md index pointed at the four project memories — OHM deployment host, Wiggleverse org structure, Open Human Model framing, ohm-rfc-app-flotilla project. The last one names this
repo's path, the SPEC as the binding contract, and the §20 slicing
plan as the menu of build sessions. The slug "slice 4" was
unambiguous against §20.4 — "Full deploy gesture + /api/health
verification, ships as v0.4.0."
SPEC §20.4 read in full:
Goal.
ohm-rfc-app-flotilla deploy ohm-rfc-appperforms the real deploy gesture against the live OHM VM, ending with successful/api/healthverification.Scope.
- Schema:
deploys,health_snapshotstables per §4.3.- CLI verbs:
deploy <deployment>,deploy abort,deploy status,deploy watch,deploy log.- Real SSH gesture via
gcloud compute ssh. Per-phase execution per §8.1 with fail-stop semantics per §8.2.- The §9 probe loop with the polling shape and verdict table from §9.2 and §9.3.
- Concurrency lock per §8.3.
- Atomic
.envwrite on the VM per §8.1 step 6.- Tests:
- Unit: phase orchestration with mocked subprocess and HTTP clients.
- Integration (run against the real OHM VM, gated by an env var): deploy the current pin (a no-op deploy), verify success.
- The first real deploy: take OHM from its current state through ohm-rfc-app-flotilla's gesture. If
ohm.wiggleverse.org/api/healthis currently misbehaving (see the session-start observation about the{"detail":"Not Found"}response), the session debugs that as part of bringing the verify loop green.Exit criteria.
ohm-rfc-app-flotilla deploy ohm-rfc-appagainst the live VM completes the 9-phase gesture, thedeploysrow recordsoutcome=succeeded, and/api/healthreports the deployed version. Ships as v0.4.0.
Cross-referenced §4.3 (the row shapes for deploys and
health_snapshots), §8.1 (the 9-phase sequence to actually run, not
just render), §8.2 (the per-phase failure table — fail-stop, no
auto-rollback, leave the VM in a state the operator can diagnose),
§8.3 (the in-flight lock semantics), §9 (the probe loop budget, verdict
table, and standalone-status forms), §10 (the deploy log shape), §11.3
(the gcloud compute ssh invocation form), and §12.1 (the full CLI
verb inventory).
Repo state at session start:
$ git log --oneline -3
491e72e Release 0.3.0: Slice 3 — pin reading + dry-run deploy
581687a Release 0.1.0: Slice 1 — registry + non-secret overlay
a9a9aaa Release 0.0.1: initial spec
$ git status
(clean)
Repo is clean from Session 0006.0's tail; nothing in flight to reconcile.
Codebase read: cli.py (375 lines, Click groups for deployment,
overlay, secret, pin, plus deploy with --dry-run-only),
db.py, registry.py (~250 lines, dataclasses + SQLite CRUD),
secrets.py (read-only Secret Manager wrapper), pin.py (Gitea
raw-file reader with certifi + explicit User-Agent), plan.py (the
dry-run plan assembler), migrations 001–003, conftest fixtures
(tmp_data_dir, conn, runner wrapping CliRunner with
expect_exit), existing tests (74 passing).
Before writing code, sketched the plan back to the user — the eight
discrete deliverables (migration, ssh.py, health.py, deploy_log.py,
deploy.py, cli.py wiring, tests, version+changelog) — then loaded the
TaskCreate/TaskUpdate tools to track them.
Turn 2 — User
It should be in the spec
Turn 2 — Assistant
Acknowledgment of "the slice plan is in the spec, work from it." No
re-asking; SPEC §20.4 had been read already, and the eight-deliverable
sketch was an enumeration of §20.4's scope, not a substitute for it.
Proceeded to create the eight tasks via TaskCreate so progress would
be visible turn-by-turn, then committed to starting with the migration
and working up the stack.
Turn 3 — User
ok go for it
Turn 3 — Assistant
Implemented the slice end-to-end. The narrative below walks the work in the order it was built — bottom-up from schema to CLI — though the tool calls in the session interleaved more, with the test files written last and triggering a couple of targeted fixes.
Migration 004
ohm_rfc_app_flotilla/migrations/004_deploys_health_snapshots.sql:
-- Slice 4: deploy log + health-probe snapshots. Per SPEC §4.3, §8, §9, §10.
--
-- `deploys` is append-only after `outcome` lands on a terminal value
-- (succeeded/failed/aborted). An `in_progress` row is the per-deployment
-- concurrency lock from §8.3.
--
-- `phases` and `post_deploy_health` are TEXT (JSON-encoded). SQLite stores
-- them as blobs; the Python layer (deploy_log.py) does the JSON dance.
--
-- `overlay_snapshot_hash` is the §8.1 step 1d hash — sha256 over the
-- canonical-JSON of {overlay values, secret-ref strings}. Never contains
-- secret bytes (§3 invariant 1).
CREATE TABLE deploys (
id INTEGER PRIMARY KEY AUTOINCREMENT,
deployment_id INTEGER NOT NULL REFERENCES deployments(id) ON DELETE CASCADE,
triggered_by TEXT NOT NULL,
triggered_at TEXT NOT NULL DEFAULT (datetime('now')),
target_version TEXT NOT NULL,
overlay_snapshot_hash TEXT NOT NULL,
phases TEXT NOT NULL DEFAULT '[]',
post_deploy_health TEXT,
outcome TEXT NOT NULL DEFAULT 'in_progress',
error_summary TEXT,
finished_at TEXT,
CHECK (outcome IN ('in_progress', 'succeeded', 'failed', 'aborted'))
);
CREATE INDEX idx_deploys_deployment ON deploys(deployment_id);
CREATE INDEX idx_deploys_outcome ON deploys(outcome);
-- §8.3: only one in_progress row per deployment at a time.
CREATE UNIQUE INDEX idx_deploys_one_in_progress
ON deploys(deployment_id)
WHERE outcome = 'in_progress';
CREATE TABLE health_snapshots (
id INTEGER PRIMARY KEY AUTOINCREMENT,
deployment_id INTEGER NOT NULL REFERENCES deployments(id) ON DELETE CASCADE,
polled_at TEXT NOT NULL DEFAULT (datetime('now')),
http_status INTEGER,
body_version TEXT,
body_status TEXT,
body_raw TEXT
);
CREATE INDEX idx_health_snapshots_deployment ON health_snapshots(deployment_id, polled_at DESC);
The partial unique index is the load-bearing piece — it makes the §8.3
"only one in-flight deploy per deployment" invariant a database fact,
not an application convention. A second INSERT … outcome='in_progress'
for the same deployment_id raises sqlite3.IntegrityError; the
Python layer catches it and renders the explicit "deploy already in
flight" error from §8.3.
Phases and post-deploy-health are TEXT columns holding JSON. SQLite
doesn't natively check JSON validity (no CHECK(json_valid(…)) in v1
— could add later if a corruption case appears, but probably unneeded
since only flotilla writes these rows). The CHECK constraint on
outcome is the only domain validator at the schema level.
finished_at is NULL while a deploy is in flight, populated at the
finalize moment. Useful for deploy log rendering ("started 12s ago")
without depending on the row's mtime.
ohm_rfc_app_flotilla/ssh.py
Three pieces: SshTarget dataclass, SshResult dataclass, run_remote
function.
@dataclass(frozen=True)
class SshTarget:
vm_name: str
zone: str
project: str
tunnel_through_iap: bool = False
@dataclass(frozen=True)
class SshResult:
returncode: int
stdout: str
stderr: str
@property
def ok(self) -> bool:
return self.returncode == 0
def run_remote(
target: SshTarget,
command: str,
*,
input: bytes | None = None,
timeout: float = 600.0,
gcloud_path: str | None = None,
) -> SshResult:
gcloud = gcloud_path or shutil.which("gcloud")
if not gcloud:
raise SshError(
"gcloud CLI not found on PATH — install Google Cloud SDK and run "
"`gcloud auth login` + `gcloud auth application-default login`"
)
argv = [
gcloud, "compute", "ssh", target.vm_name,
"--zone", target.zone,
"--project", target.project,
"--quiet",
]
if target.tunnel_through_iap:
argv.append("--tunnel-through-iap")
argv.extend(["--command", command])
try:
proc = subprocess.run(
argv, input=input, capture_output=True, timeout=timeout,
)
except subprocess.TimeoutExpired as e:
raise SshError(f"ssh to {target.vm_name} timed out after {timeout}s …") from e
except FileNotFoundError as e:
raise SshError(f"gcloud not executable: {e}") from e
return SshResult(
returncode=proc.returncode,
stdout=_decode(proc.stdout),
stderr=_decode(proc.stderr),
)
Three design choices worth recording:
- Non-zero exit codes return
SshResult, do not raise. The §8.2 per-phase failure table sometimes records state without bailing — the preflightsystemctl is-activeprobe is record-only, for instance. Letting the caller examinereturncode+stderrkeeps phase code straight-line. Only subprocess-level failure (timeout, gcloud not found) raisesSshError. input: bytes | Noneparameter pipes to stdin. Used by §8.1 step 6's atomic.envwrite to deliver the env body without landing it on the operator's disk. The bytearray (zeroed after the call) → subprocess stdin → SSH stream → remotecat > tmp. Single-hop, scrubbable end-to-end.tunnel_through_iapreads from the deployment record. Per SPEC §11.3 and the OHM-side §13.2 (target_vm_tunnel_through_iap: falsefor current OHM, because port 22 is open by legacy firewall). Future "harden SSH via IAP-only" §19.2 candidate tightens this default.
The whole module is dependency-injected: deploy.py accepts an
ssh_runner: Callable[..., SshResult] parameter that defaults to
ssh.run_remote. Unit tests pass a _SshRecorder that records
commands and returns scripted results, never touching gcloud.
ohm_rfc_app_flotilla/health.py
The §9 probe — one-shot probe_once, the evaluate(reading, target)
function that applies the §9.3 verdict table, and the probe_loop
that runs the §9.2 polling loop.
class Verdict(str, Enum):
SUCCESS = "success"
MISMATCH = "mismatch"
DEGRADED = "degraded"
ENDPOINT_MISSING = "endpoint_missing"
TRANSIENT = "transient"
@dataclass(frozen=True)
class HealthReading:
polled_at: str
http_status: int | None
body_version: str | None
body_status: str | None
body_raw: str | None
error: str | None = None
def evaluate(reading: HealthReading, target_version: str) -> Verdict:
if reading.http_status is None:
return Verdict.TRANSIENT # connection refused, DNS, TLS
code = reading.http_status
if code == 200:
if reading.body_version == target_version and reading.body_status == "ok":
return Verdict.SUCCESS
return Verdict.MISMATCH
if code == 503 and reading.body_status == "degraded":
return Verdict.DEGRADED
if 400 <= code < 500:
return Verdict.ENDPOINT_MISSING
if 500 <= code < 600:
return Verdict.TRANSIENT
return Verdict.TRANSIENT
probe_once wraps urllib.request.urlopen and never raises. Every
failure path produces a HealthReading — HTTP error responses become
readings with the failing status code; URL/network errors become
readings with http_status=None and an error field. This means the
loop logic is uniformly "always get a reading, always evaluate, decide
what to do."
probe_loop implements the §9.2 polling shape:
def probe_loop(
url: str,
target_version: str,
*,
prober: Prober = probe_once,
on_reading: Callable[[HealthReading, Verdict], None] | None = None,
initial_interval: float = 2.0,
max_interval: float = 30.0,
budget_seconds: float = 300.0,
per_request_timeout: float = 10.0,
sleep: Callable[[float], None] = time.sleep,
monotonic: Callable[[], float] = time.monotonic,
) -> LoopResult:
start = monotonic()
deadline = start + budget_seconds
interval = initial_interval
attempts = 0
last_reading, last_verdict = None, None
while True:
attempts += 1
reading = prober(url, per_request_timeout)
verdict = evaluate(reading, target_version)
last_reading, last_verdict = reading, verdict
if on_reading is not None:
on_reading(reading, verdict)
if verdict == Verdict.SUCCESS:
break
if verdict == Verdict.ENDPOINT_MISSING:
break
if monotonic() >= deadline:
break
if attempts >= 5:
interval = min(interval * 2, max_interval)
remaining = max(0.0, deadline - monotonic())
if remaining <= 0:
break
sleep(min(interval, remaining))
duration_ms = int((monotonic() - start) * 1000)
return LoopResult(verdict=last_verdict, final_reading=last_reading,
attempts=attempts, total_duration_ms=duration_ms)
The §9.2 budget shape baked in:
- Initial 2s interval, exponential after 5 transient attempts, capped at 30s (§9.2 says "doubles each subsequent attempt to a maximum of 30 seconds").
- 300s (5min) total budget. Sleep caps to remaining budget so the loop doesn't overshoot.
- 10s per-request timeout (connect + read combined; urllib's
timeoutis conservative here). - SUCCESS and ENDPOINT_MISSING exit immediately; MISMATCH and DEGRADED keep polling for budget (§9.4: "a slow startup may still land"); TRANSIENT keeps polling.
sleep and monotonic are dependency-injected — the unit tests use
a _FakeClock that records sleeps without actually sleeping, so a 5-
minute test budget runs in microseconds.
on_reading is the per-reading callback the deploy gesture uses to
record each reading into the health_snapshots table — every probe
counts, not just the final one. The deploy watch CLI verb also
plugs into this callback to log every reading it sees.
The TLS+UA fixups from Slice 3's pin.py carry over: _ssl_context()
uses certifi.where() if available (transitively present via
google-cloud-secret-manager), and the User-Agent is the explicit
ohm-rfc-app-flotilla (health prober) string so Cloudflare-fronted
ohm.wiggleverse.org doesn't 403 the prober.
ohm_rfc_app_flotilla/deploy_log.py
The persistence layer for deploys and health_snapshots. Two
dataclasses (DeployRow, HealthSnapshotRow) plus a PhaseRecord
for the per-phase JSON entries, then a handful of operations:
open_deploy(conn, name, *, triggered_by, target_version, overlay_snapshot_hash) → int— INSERT a row withoutcome='in_progress', return its id. CatchesIntegrityErrorand translates toConcurrencyError(the §8.3 lock).append_phase(conn, deploy_id, PhaseRecord)— append to thephasesJSON list. Done one phase at a time, so a row's progress is observable mid-flight viadeploy log <id>.finalize_deploy(conn, deploy_id, *, outcome, error_summary, post_deploy_health)— set the terminal outcome (succeeded/failed/aborted), setfinished_at, persistpost_deploy_healthif present. Releases the §8.3 lock.abort_in_flight(conn, name) → DeployRow— find the in-flight row, mark it aborted, return it. RaisesNotFoundErrorif nothing is in flight.get_deploy,list_deploys,get_in_flight— readers.record_health_snapshot,list_health_snapshots— thehealth_snapshotstable.redact(text, secret_bytes) → str— the §10.2 stderr scrubber.
The redact shape:
def redact(text: str, secret_bytes: Iterable[bytes]) -> str:
if not text:
return text
redacted = text
for raw in secret_bytes:
if not raw:
continue
try:
needle = raw.decode("utf-8", errors="replace")
except Exception:
continue
if needle and needle in redacted:
redacted = redacted.replace(needle, "[REDACTED]")
return redacted
Cheap substring scan. Not cryptographically robust, but the secret
bytes in question are random tokens — a substring match against the
exact resolved payload is conclusive. The redacted string is the
error_summary column value; the unredacted stderr never lands in
flotilla state.
The ConcurrencyError translation:
def open_deploy(conn, deployment_name, *, ...) -> int:
dep = get_deployment(conn, deployment_name)
try:
cur = conn.execute(
"INSERT INTO deploys (deployment_id, …, outcome) "
"VALUES (…, 'in_progress')",
(dep.id, …),
)
except sqlite3.IntegrityError as e:
in_flight = conn.execute(
"SELECT id FROM deploys WHERE deployment_id = ? AND outcome = 'in_progress'",
(dep.id,),
).fetchone()
raise ConcurrencyError(
f"deploy already in flight for {deployment_name!r} "
f"(deploys.id={in_flight['id'] if in_flight else '?'}); "
"`flotilla deploy abort` to release the lock"
) from e
return int(cur.lastrowid)
Catches a generic IntegrityError, then looks up the in-flight row's
id to put it in the error message. The CLI catches ConcurrencyError
and prints the explicit message, including the existing deploys.id
so the operator can investigate before aborting.
ohm_rfc_app_flotilla/deploy.py
The 9-phase orchestration. The longest file in the slice — every step
in §8.1 gets a phase body with its own SSH command, success/failure
discrimination, and _PhaseFailure raise on non-ok outcome. The
_PhaseRunner helper records start/end timestamps and the phase
record into the deploys row.
Three structural pieces:
_ResolvedSecrets — the bytearray context for §3 invariant 1
- §8.1 step 1c "bytes held in memory for the duration of this command and zeroed at the end":
class _ResolvedSecrets:
def __init__(self) -> None:
self._values: dict[str, bytearray] = {}
def __enter__(self): return self
def __exit__(self, *exc) -> None: self.scrub()
def add(self, env_key: str, payload: bytes) -> None:
self._values[env_key] = bytearray(payload)
def as_bytes_dict(self) -> dict[str, bytes]:
return {k: bytes(v) for k, v in self._values.items()}
def scrub(self) -> None:
for ba in self._values.values():
for i in range(len(ba)):
ba[i] = 0
self._values.clear()
run_deploy uses it via try / finally around the entire orchestration
body so an exception unwind still scrubs. Python's immutable strings
mean we can't get perfect scrubbing — the GC will see a bytes
intermediate if anything calls .decode() — but the bytearray stays
in our control end-to-end and the byte pattern is overwritten before
the bytearray drops out of scope.
compose_env_body(overlay, resolved_secrets) → bytearray —
renders the dotenv body from the overlay dict + resolved secret
bytes:
def compose_env_body(overlay, resolved_secrets) -> bytearray:
body = bytearray()
combined = []
for k, v in overlay.items():
combined.append((k, v))
for k, raw in resolved_secrets.items():
combined.append((k, raw.decode("utf-8", errors="surrogateescape")))
combined.sort(key=lambda pair: pair[0])
for k, v in combined:
line = f"{k}={_escape_env_value(v)}\n"
body.extend(line.encode("utf-8", errors="surrogateescape"))
return body
def _escape_env_value(v: str) -> str:
escaped = (
v.replace("\\", "\\\\")
.replace("\"", "\\\"")
.replace("\n", "\\n")
.replace("\r", "\\r")
)
return f'"{escaped}"'
Always-quoted double-quoted form with the standard dotenv escape
sequences. Sorted by key — not specified by the spec, but a
deterministic order helps the operator diff between deploys, and
the framework doesn't care about order. surrogateescape handles
the (unlikely) case of a secret payload that isn't valid UTF-8 —
the byte payload round-trips exactly.
The 9-phase orchestration lives in run_deploy. Step 1's
validation (resolve pin, resolve secrets, compute hash) happens
before the deploys row opens — a pin or secret resolution
failure raises DeployError without polluting state. The
open_deploy call is the moment the §8.3 lock latches. From that
moment, every phase runs through _PhaseRunner.run which records
the phase record, and any phase failure goes through
_finalize_failure which writes outcome=failed plus the redacted
error summary.
Each phase body is a small closure passed to runner.run:
def _phase2() -> tuple[str, str | None]:
r = ssh_runner(target, f"test -d {shlex.quote(install)}")
if r.returncode != 0:
raise _PhaseFailure(
f"install dir {install} missing on {dep.target_vm_name} "
"(fresh-VM provisioning is out of v1)",
stderr=r.stderr,
)
r2 = ssh_runner(target, f"systemctl is-active {shlex.quote(unit)} || true")
pre_state["unit_active"] = r2.stdout.strip() or "unknown"
return f"install ok; pre-deploy is-active={pre_state['unit_active']}", None
ok, _, _ = runner.run(2, "preflight", _phase2)
if not ok:
return _finalize_failure(conn, deploy_id, "preflight", …)
Pattern repeats for each phase. The fail-stop discipline (§8.2) is
"if not ok, immediately return through _finalize_failure" — every
phase that fails halts the gesture and records the failing step.
Phase 6 — the atomic .env write. This is the load-bearing piece
worth quoting in full:
def _phase6() -> tuple[str, str | None]:
env_path = install + "/backend/.env"
tmp = env_path + ".new"
inner = (
"set -e\n"
"umask 077\n"
f"cat > {shlex.quote(tmp)}\n"
f"mv -f {shlex.quote(tmp)} {shlex.quote(env_path)}\n"
)
cmd = f"sudo -u {shlex.quote(user)} sh -c {shlex.quote(inner)}"
r = ssh_runner(target, cmd, input=bytes(env_body), timeout=120.0)
if r.returncode != 0:
raise _PhaseFailure(f".env write failed: …", stderr=r.stderr)
return f"wrote {env_path} ({len(env_body)} bytes, mode 0600)", None
env_body = compose_env_body(overlay_dict, resolved.as_bytes_dict())
try:
ok, _, _ = runner.run(6, "write .env", _phase6)
finally:
# Zero the env body bytes regardless of outcome (§8.1 step 6d).
for i in range(len(env_body)):
env_body[i] = 0
set -e so cat failure halts before mv. umask 077 makes the
temp file mode 0600. sudo -u <user> makes it owned by the service
user. Both cat's target and mv's destination live in the same
directory, so the rename is rename(2) — atomic on the local
filesystem. Stdin pipes from bytes(env_body); subprocess delivers
it to the gcloud ssh stream, which delivers it to remote cat. The
env body never lands on the operator's disk. The finally zeroes
the bytearray after the call returns regardless of success or
failure.
Phase 7 — restart and wait. Done on the VM side rather than client-side polling, to avoid 30 round-trip costs:
inner = (
f"systemctl restart {shlex.quote(unit)} && "
f"for i in $(seq 1 {int(restart_wait_seconds)}); do "
f" systemctl is-active --quiet {shlex.quote(unit)} && exit 0; "
" sleep 1; "
"done; "
"echo 'is-active timed out' >&2; exit 1"
)
cmd = f"sudo sh -c {shlex.quote(inner)}"
30 iterations, 1s sleep each — matches §8.1 step 7b's "up to 30 seconds" budget. The shell exits as soon as the unit reports active, or fails after the budget.
Phase 8 — verify /api/health. Calls into health.probe_loop with
an on_reading callback that persists every reading into
health_snapshots:
def _phase8() -> tuple[str, str | None]:
nonlocal loop_result
def _on_reading(reading, verdict):
deploy_log.record_health_snapshot(
conn, deployment_name,
http_status=reading.http_status,
body_version=reading.body_version,
body_status=reading.body_status,
body_raw=reading.body_raw,
)
loop_result = probe_loop(
dep.health_url, target_version,
prober=prober, on_reading=_on_reading,
initial_interval=health_initial_interval,
budget_seconds=health_budget_seconds,
)
if loop_result.verdict == Verdict.SUCCESS:
return (f"/api/health reports v{loop_result.final_reading.body_version} …", None)
raise _PhaseFailure(_verdict_failure_detail(loop_result, target_version, dep), stderr=None)
_verdict_failure_detail produces the §9.4-style explicit message
for each non-SUCCESS verdict. Example, MISMATCH:
/api/health reports version '0.2.2', expected '0.3.0' (status='ok'). The systemd restart may not have picked up the new code.
(The spec example in §9.4 also includes a recovery hint with the
explicit gcloud compute ssh invocation; in this implementation the
hint stays in the spec rather than the runtime message, since the
deployment record carries the VM coordinates and the CLI can
synthesize the hint as a separate output line if Slice 5 finds it
useful.)
resolve_operator_account() — shells out to gcloud config get-value account to learn the operator's ADC identity for
triggered_by. Falls back to "unknown" if gcloud is absent or the
call fails. Tests pass an explicit operator= to skip the shell-out.
ohm_rfc_app_flotilla/cli.py — the deploy surface
The §12.1 inventory needed real deploy <name> (the gesture) plus
deploy abort/status/watch/log. The naive Click pattern:
@main.group("deploy", invoke_without_command=True)
@click.argument("deployment", required=False)
@click.option("--dry-run", is_flag=True, default=False)
@click.option("--json", "as_json", is_flag=True)
@click.pass_context
def deploy_cmd(ctx, deployment, dry_run, as_json):
if ctx.invoked_subcommand is not None:
return
…
Looked correct, tested wrong:
$ flotilla deploy ohm-rfc-app --dry-run
Usage: main deploy [OPTIONS] [DEPLOYMENT] COMMAND [ARGS]...
Try 'main deploy --help' for help.
Error: No such command '--dry-run'.
Click parses the group's positional [DEPLOYMENT] greedily — it
consumes ohm-rfc-app as DEPLOYMENT, then treats the rest (--dry-run)
as the subcommand name. Failure mode is structural: a group with both
a positional argument and invoke_without_command=True can't
distinguish "args for the implicit verb" from "args for a subcommand."
Resolution: a custom Group subclass that routes any non-subcommand
first argument to a hidden run subcommand:
class _DeployGroup(click.Group):
"""`deploy` is both a group (abort/status/watch/log) and a verb
(`deploy <name>`). When the first arg isn't a known subcommand and
isn't an option, route it to the hidden `run` subcommand so the
SPEC §12.1 surface works alongside `flotilla deploy abort <name>`."""
def resolve_command(self, ctx, args):
cmd_name = args[0] if args else ""
if cmd_name and cmd_name not in self.commands and not cmd_name.startswith("-"):
args = ["run"] + list(args)
return super().resolve_command(ctx, args)
@main.group("deploy", cls=_DeployGroup)
def deploy_cmd():
"""Deploy a deployment or operate on its deploy log / health probe."""
@deploy_cmd.command("run", hidden=True)
@click.argument("deployment")
@click.option("--dry-run", is_flag=True, default=False)
@click.option("--json", "as_json", is_flag=True)
def deploy_run(deployment, dry_run, as_json):
if dry_run:
_do_dry_run(deployment, as_json)
return
if as_json:
_exit_with_error("--json is only meaningful with --dry-run on `deploy`")
_do_real_deploy(deployment)
flotilla deploy ohm-rfc-app --dry-run → resolve_command sees
ohm-rfc-app is not in {abort, status, watch, log, run} and doesn't
start with -, prepends run → Click invokes run ohm-rfc-app --dry-run. flotilla deploy abort ohm-rfc-app → abort is in
commands, no rewriting → Click invokes abort ohm-rfc-app. The §12.1
surface is preserved; the implementation seam stays in the CLI module.
run is hidden=True so it doesn't appear in deploy --help —
operators never type "run" explicitly. The help output for flotilla deploy --help reads cleanly:
Usage: ohm-rfc-app-flotilla deploy [OPTIONS] COMMAND [ARGS]...
Deploy a deployment or operate on its deploy log / health probe.
Commands:
abort Mark the in-flight deploy for DEPLOYMENT as aborted …
log List recent deploys, or show one deploy in full.
status One-shot /api/health probe; records the reading and renders it.
watch Periodic /api/health probe.
The five real subcommands:
deploy abort <name>— callsdeploy_log.abort_in_flight.deploy status <name>— one-shothealth.probe_once, persists viadeploy_log.record_health_snapshot, prints the line.--jsonfor scripting.deploy watch <name>— periodic probe loop driven bytime.sleep, persists every reading, prints each, exits on Ctrl-C.deploy log <name>— list recent deploys (default limit 20).deploy log <name> <id>— render one deploy in full, including the phases array and the post-deploy-health JSON.
pin check also got an upgrade — Slice 3 had it print a "Slice 4
will extend this to compare against the last successful deploy"
forward reference. Slice 4 replaced that:
@pin.command("check")
@click.argument("deployment")
def pin_check(deployment):
…
succeeded = [d for d in deploy_log.list_deploys(conn, dep.name, limit=50)
if d.outcome == "succeeded"]
if not succeeded:
click.echo("(no successful deploys yet — `flotilla deploy ohm-rfc-app` to land one)")
return
last = succeeded[0]
if last.target_version == version:
click.echo(f"pin matches last deploy (v{last.target_version}, deploys.id={last.id})")
else:
click.echo(
f"pin AHEAD of last deploy: pin={version}, "
f"last deployed=v{last.target_version} (deploys.id={last.id})"
)
Three states: no deploys yet, pin matches last deploy, pin advanced
past last deploy. The "AHEAD" case is the one §20.5's bring-up flow
("echo 0.3.1 > ohm-rfc/.rfc-app-version && git commit && git push; flotilla pin check ohm-rfc-app") wants to see — the pin moved, but
the deploy hasn't happened yet.
Test suite
Five new test files, 58 new tests (74 → 132 total):
-
tests/test_health.py— 12 tests. Each row of the §9.3 verdict table by case (success/mismatch/degraded/endpoint_missing/transient/ connection-error). Probe-loop happy path (immediate success, no sleeps), transient retry then success, endpoint-missing exits immediately, mismatch keeps polling for budget, backoff caps at max_interval. The_FakeClockfixture recordssleep()calls and advancesmonotonic()accordingly — a 5-minute budget test runs in microseconds. -
tests/test_ssh.py— 6 tests. Argv assembly (gcloud path, zone/project flags, command escaping),--tunnel-through-iapinclusion when set, stdin pipe round-trip, non-zero return without raising, timeout raisesSshError, missing gcloud raisesSshError. All tests monkeypatchshutil.whichandsubprocess.run— no real gcloud invocation. -
tests/test_deploy_log.py— 14 tests.open_deployround-trip, the §8.3 lock (second open while one is in flight raisesConcurrencyError), finalize releases the lock, abort path,append_phaseround-trip, finalize recordspost_deploy_health, invalid outcome rejected, list returns reverse-chronological,get_in_flightreturns None when idle,record_health_snapshot,redacthappy path and edge cases. -
tests/test_deploy.py— 12 tests. The whole-orchestration surface. Happy-path run records 9 succeeded phases with post_deploy_health captured. Phase-2 missing-install-dir fails fast at preflight with the correct error_summary. Pin-error raisesDeployErrorbefore anydeploysrow opens. Unreadable secret raisesDeployErrorbefore anydeploysrow opens. §8.3 lock raisesDeployErrorwhen an in-flight row is pre-seeded. Env body composition: sort order, quote/backslash/newline escaping. The_ResolvedSecrets.scrub()zeros the bytes.resolve_operator_accountfalls back to "unknown" when gcloud is absent.Two tests deserve specific note:
Secret-byte leak protection. A phase that simulates the SSH command failing and dumping the stub secret into stderr:
class LeakyFail: def __call__(self, target, command, *, input=None, timeout=None, **kw): if "git fetch" in command: err = f"fatal: token={STUB_SECRET_BYTES.decode()} was rejected" return _fail(stderr=err) return _ok() outcome = deploymod.run_deploy(…) assert outcome.outcome == "failed" assert STUB_SECRET_BYTES.decode("ascii") not in (outcome.error_summary or "") assert "[REDACTED]" in (outcome.error_summary or "")The redaction holds — the recorded error_summary has
[REDACTED]where the stub bytes would have appeared.Env-body stdin pipe. A test that lets the whole orchestration succeed and then inspects the recorded SSH calls for the .env write — the bytes appear in the
inputparameter (the stdin pipe to ssh) but never in any recordeddeploysrow or serialized output:env_writes = [c for c in ssh_rec.calls if c["input"] is not None] assert len(env_writes) == 1 body = env_writes[0]["input"] assert b"GITEA_URL=" in body assert b"SECRET_KEY=" in body assert STUB_SECRET_BYTES in body # bytes flow through into VM-bound stdin row = deploy_log.get_deploy(conn, "ohm-rfc-app", outcome.deploy_id) flat = json.dumps(row.to_dict()) assert STUB_SECRET_BYTES.decode("ascii") not in flat -
tests/test_cli_deploy.py— 14 tests. CLI smoke for every new verb. Realdeployhappy path (mocks every external call — pin, secret, ssh, prober — and asserts "ohm: deployed v0.3.0"), unknown deployment fails nonzero,--jsonwithout--dry-runfails with the explicit message,--dry-runstill works (Slice 3's surface preserved), abort releases lock, abort with nothing in flight fails, status renders human + JSON forms, status records a snapshot, log empty / list / full-render / JSON / unknown id, pin check now shows "matches last deploy" or "AHEAD" given the relevant deploy-log state.
The integration test the §20.4 scope mentions ("run against the
real OHM VM, gated by an env var") was not shipped this session.
Reasoning: a test that shells out to live gcloud would only work on
a workstation already configured against wiggleverse-ohm plus the
operator's ADC — that configuration is the bring-up the live deploy
session performs, not something to bake into CI. The 132 unit tests
cover the orchestration shape and every phase boundary; the live
gesture is the §20.4 final exit criterion and the operator gesture
this release is gated on. v0.5.0/§20.5 decides whether to add a
gated integration test as part of hardening.
Version + changelog
VERSION bumped from 0.3.0 to 0.4.0. pyproject.toml's
version = "0.3.0" → "0.4.0". pip install -e . re-run to pick
up the entry-point version; ohm-rfc-app-flotilla --version now
reports 0.4.0.
CHANGELOG.md got a 0.4.0 — 2026-05-27 entry above the existing
0.3.0 section. The entry covers every shipped piece — migration
004 + the partial-unique-index for §8.3, the four new modules, the
CLI extensions, the Click custom-group pattern, the pin check
upgrade, the test split, the integration-test deferral, the spec
status (no §19.3 corrections this slice). Closes with a pre-1.0
RFC-2119 upgrade-steps block:
Upgrade steps:
- An operator upgrading from v0.3.0 MUST run any flotilla CLI command once against an existing SQLite to apply migration
004_deploys_health_snapshots.sql. The migration is additive (two new tables, no changes to existing rows) and theopen_db()migration runner picks it up automatically on the next invocation.- No data-format changes to
deployments,overlay_entries,secret_refs, orschema_migrations.
Final test run
$ .venv/bin/python -m pytest 2>&1 | tail -5
tests/test_secrets.py .................... [ 95%]
tests/test_ssh.py ...... [100%]
============================= 132 passed in 0.20s ==============================
132 tests, all green. Repo state at end of session:
$ git status
M CHANGELOG.md, VERSION, pyproject.toml,
ohm_rfc_app_flotilla/cli.py,
tests/test_plan.py
?? ohm_rfc_app_flotilla/deploy.py,
ohm_rfc_app_flotilla/deploy_log.py,
ohm_rfc_app_flotilla/health.py,
ohm_rfc_app_flotilla/migrations/004_deploys_health_snapshots.sql,
ohm_rfc_app_flotilla/ssh.py,
tests/test_cli_deploy.py,
tests/test_deploy.py,
tests/test_deploy_log.py,
tests/test_health.py,
tests/test_ssh.py
— commit + tag pending; the user's typical pattern is a single
release commit per slice (matching 491e72e Release 0.3.0: …), and
the live first-deploy gesture (§20.4's "first real deploy" line) is
the operator gesture between this session and the v0.5.0 hardening
session.
End-of-turn summary delivered to the user noting:
- 9-phase gesture, §9 probe loop, deploy-log + health-snapshots schema, CLI verbs, Click custom-group, secret-byte scrubbing.
- 58 new tests, 132 total, all passing.
- VERSION + CHANGELOG bumped to 0.4.0.
- Live first deploy against the OHM VM is the remaining operator
gesture; can't be run from here without
gcloud auth application-default loginagainstwiggleverse-ohm.
Turn 4 — User
ok, write out a full transcript (EVERYTHING) to ~/git/ohm-infra
Turn 4 — Assistant
This document.
Hand-off for Slice 5
For the operator + the Slice 5 session.
Slice 5's brief (SPEC §20.5): Use flotilla to drive a real
rfc-app version increment end-to-end. Bump
ohm-rfc/.rfc-app-version to whatever has accumulated since
v0.3.0 on the framework side (if nothing has, cut a 0.3.1 patch
on rfc-app specifically to exercise the path); run flotilla pin check ohm-rfc-app (verify the AHEAD render); run flotilla deploy ohm-rfc-app; verify /api/health reports the new
version. Tighten failure paths discovered in Slice 4. Add --json
on every read verb per §12.3. Write docs/operator-guide.md for
laptop bring-up. Ships as v1.0.0.
Before Slice 5 runs, the operator gesture between sessions:
- Commit + tag v0.4.0. Single release commit, matching the
491e72e Release 0.3.0: Slice 3 — pin reading + dry-run deployshape. The working tree is exactly the v0.4.0 surface; no reconciliation needed. - Live first deploy (§20.4's exit criterion). Roughly:
If
$ gcloud auth login $ gcloud auth application-default login $ flotilla deployment list $ flotilla pin check ohm-rfc-app $ flotilla deploy ohm-rfc-app/api/healthis currently returning{"detail":"Not Found"}(the session-start observation noted in SPEC §20.4 scope), the live deploy is the moment to debug it. Likely candidates: the/api/healthroute is mounted under an/api/prefix that isn't being routed through nginx, or the framework version on the VM is somehow pre-0.2.3 despite the pin claiming 0.3.0. The §9.3 ENDPOINT_MISSING verdict produces the exact error message: "endpoint not reachable — is the framework version pre-0.2.3?"
Findings to pre-load for Slice 5:
-
The
_DeployGroupClick pattern is the working example for any future "group that's also a verb" surface. If Slice 5 addspin set(a §19.2 candidate flagged in Slice 3), the same_PinGroupshape can support it. Today'spingroup hasshowandchecksubcommands; addingpin <name>as a default (e.g., for "show the current pin") would use the sameresolve_commandrewrite. -
The DI-friendly orchestration in
deploy.pyis the template for any future external-call gesture.pin_reader,secret_reader,ssh_runner,proberare all injectable; the tests never touch the real gcloud or network. Any new gesture in Slice 5 (e.g., a rollback verb) should follow the same shape. -
The §10.2
redact()helper takes an iterable of byte payloads. If Slice 5 adds new sources of sensitive bytes (e.g., a Gitea token for private-corpus support), thread them through the same redaction call soerror_summarystays clean. -
flotilla deploy log <name>'s rendering is intentionally spare in v0.4.0. The phases list is one line per phase plus optional detail. Slice 5's hardening pass might want a richer rendering — total duration, per-phase duration computed from start/end timestamps, a separate "failed phase" callout. -
The §8.3 lock is best-tested by pre-seeding an in_progress row rather than trying to run two real deploys concurrently (which the SQLite serialization makes hard to observe anyway). See
test_run_deploy_blocks_when_in_flight. -
The health probe's
Verdict.TRANSIENTonhttp_status=Noneis load-bearing. A network outage during the probe loop should not bail the verify phase — connection refused is normal during the first few seconds aftersystemctl restart. Don't tighten this without thinking about the §8.1 step 7 → 8 boundary. -
The bytearray scrubbing isn't cryptographically perfect.
bytes.decode()produces an immutable string the GC sees. The intermediate exists for thesurrogateescaperound-trip incompose_env_bodyand for the env-body bytes that go into the subprocess. The bytearray itself is scrubbed; the transient immutable copies are not. Probably good enough — the secret bytes live for the duration of one process invocation, which exits seconds after the deploy completes — but record the limitation here so Slice 5 doesn't pretend otherwise. -
The migration runner is idempotent. Slice 5 can land a migration 005 (if any schema change is needed for the §20.5 hardening) without special handling; the existing
db.run_migrationspicks up new files via glob and skips ones recorded inschema_migrations. -
The auto-memory unchanged. No new memory files written this session. The slice is fully recoverable from the commit (when it lands) + CHANGELOG + this transcript; nothing about it needs to persist as cross-session context. If Slice 5 surfaces a non-obvious operator preference or a load-bearing decision not captured in the spec/code, that's when to add memory.
-
CI is still Linux + stubbed external calls. The 132 tests pass on the macOS laptop with the framework Python the operator uses; they should also pass on Linux CI without certifi (the
_ssl_context()fallback tossl.create_default_context()handles the platform difference). No network is touched by any unit test. -
/api/healthdebugging recipe, if Slice 5's live first deploy surfaces the §20.4 "currently misbehaving" observation:# what does the VM say is running? gcloud compute ssh ohm-app --zone=us-central1-a -- \ 'sudo systemctl status ohm-app && \ sudo cat /opt/ohm-app/backend/VERSION 2>/dev/null && \ cd /opt/ohm-app && git describe --tags' # what does the load balancer say? curl -i https://ohm.wiggleverse.org/api/health curl -i https://ohm.wiggleverse.org/api/healthz # alternate path some apps use # is it an nginx routing problem? gcloud compute ssh ohm-app --zone=us-central1-a -- \ 'sudo cat /etc/nginx/sites-enabled/* | grep -A 10 location'If the framework's
VERSIONfile says 0.3.0 but the endpoint 404s, it's an nginx routing issue. If the framework'sVERSIONfile says 0.2.2 or earlier, the deploy never landed correctly (the operator may have been doing manual deploys before flotilla existed). Either case, the §9.3 ENDPOINT_MISSING message gets the operator the right diagnosis line.