Files
session-history/0007/SESSION-0007.0-TRANSCRIPT-2026-05-27T19-30--2026-05-27T19-45.md
T
Ben Stull ad01c0854a #30: per-session folder layout + README + sessions.json
Restructures the flat repo into one folder per session, with a top-level
README.md (the "about sessions" page, also rendered at /docs/sessions/about
by rfc-app v0.19.0+) and a sessions.json title manifest keyed by 4-digit
NNNN. Filenames inside each folder retain the full SESSION-NNNN.M-TRANSCRIPT-…
form (no trim of the SESSION- prefix). Legacy letter-form transcripts, if
any are still at the root, are not moved — the folder convention applies
to the numeric form only.

Session 0017.0; see SESSION-PROTOCOL.md §1 amendment for the binding shape.
2026-05-28 09:12:08 -07:00

1197 lines
49 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Session G — Transcript
> Date: 2026-05-27
> Goal: Execute Slice 4 of the `ohm-rfc-app-flotilla` SPEC §20 slicing
> plan — full deploy gesture + `/api/health` verification. Per §20.4:
> add the `deploys` and `health_snapshots` tables; ship the real SSH
> gesture via `gcloud compute ssh` per §8.1 with fail-stop semantics
> per §8.2; the §9 probe loop with the §9.2 polling shape and the §9.3
> verdict table; the §8.3 concurrency lock; the §8.1 step 6 atomic
> `.env` write on the VM; new CLI verbs `deploy <name>`,
> `deploy abort`, `deploy status`, `deploy watch`, `deploy log`. Exit
> criterion: the live gesture records `outcome=succeeded` and
> `/api/health` reports the deployed version, ships as v0.4.0.
>
> Outcome: **Slice 4 code shipped at v0.4.0** on the working tree —
> not yet committed, not yet tagged, and the live first-deploy gesture
> against the OHM VM (§20.4 final exit criterion) is the operator
> gesture this release is gated on. All 132 tests pass locally (74
> baseline → 132, +58 for Slice 4). New surfaces: migration
> `004_deploys_health_snapshots.sql` (with a SQLite partial-unique-
> index implementing the §8.3 in-flight lock); `ohm_rfc_app_flotilla/
> ssh.py` (a thin DI-friendly `gcloud compute ssh` wrapper);
> `ohm_rfc_app_flotilla/health.py` (one-shot probe + §9 probe loop
> with the §9.3 verdict table verbatim); `ohm_rfc_app_flotilla/
> deploy_log.py` (CRUD for `deploys` and `health_snapshots`, phase
> appender, abort path, `redact()` for §10.2); `ohm_rfc_app_flotilla/
> deploy.py` (the 9-phase orchestration with secret-byte scrubbing
> via a `_ResolvedSecrets` bytearray context). CLI surface extended
> to the full §12.1 verb inventory. `pin check` now compares against
> the most recent successful deploy.
>
> Five findings worth flagging permanently:
>
> 1. **Click `invoke_without_command=True` does not compose with a
> positional `[DEPLOYMENT]` argument on the group.** Slice 4 needed
> `deploy <name>` (a verb) and `deploy abort <name>` (a subcommand)
> to coexist per SPEC §12.1. The naive pattern (`@main.group("deploy",
> invoke_without_command=True)` + `@click.argument("deployment",
> required=False)`) parses `deploy ohm-rfc-app --dry-run` as
> DEPLOYMENT=ohm-rfc-app, then tries to invoke a subcommand named
> `--dry-run` and fails. Resolution: a custom `_DeployGroup(click.Group)`
> subclass overrides `resolve_command` to route any non-subcommand
> first arg to a hidden `run` subcommand. The §12.1 surface is
> preserved without compromise. Implementation note now sits in
> `cli.py` so a future maintainer doesn't re-discover it.
> 2. **SQLite partial unique indexes are the right tool for the §8.3
> in-flight lock.** A `CREATE UNIQUE INDEX … WHERE outcome =
> 'in_progress'` over `deploys(deployment_id)` makes the database
> refuse a second in-flight row directly; `deploy_log.open_deploy`
> catches `sqlite3.IntegrityError` and translates it into
> `ConcurrencyError` with the explicit "deploy already in flight"
> message. No advisory locking, no application-level mutex; the
> schema enforces the invariant.
> 3. **The `/api/health` §9.3 verdict table maps cleanly to a
> `Verdict` enum + a flat `evaluate(reading, target_version)`
> function.** Each row of the table becomes one branch; the probe
> loop reads the return value and decides whether to exit
> immediately (SUCCESS, ENDPOINT_MISSING), keep polling within
> budget (MISMATCH, DEGRADED), or back off and retry (TRANSIENT).
> Connection-failure readings (`http_status is None`) fold into
> TRANSIENT — matches §9.3 row 4 "connection refused → transient".
> 4. **Secret-byte scrubbing is best done with a context-managed
> `bytearray` container.** `_ResolvedSecrets` in `deploy.py` holds
> each resolved payload as a `bytearray`; `__exit__` (always
> called, including during exception unwind) overwrites every byte
> with `0x00` and clears the dict. The env body itself is also a
> `bytearray` zeroed in a `finally` after the VM-side write
> completes. Python's immutable `str` makes truly-secure scrubbing
> impossible, but the bytearray path stays in our control end-to-
> end and the GC never sees the payload as a managed object.
> 5. **Atomic `.env` write via single SSH command piping body to
> `cat > tmp && mv -f tmp final`.** Both files live in the same
> directory, so the rename is `rename(2)` — atomic on the local
> filesystem. `umask 077` inside the `sh -c` body makes the temp
> file mode 0600; `sudo -u <user>` makes it owned by the service
> user. The stdin pipe means the env body never lands on the
> operator's disk; one bytearray, one network hop, scrubbed at
> scope exit.
>
> No SPEC corrections this slice. The implementation discoveries above
> are Click and SQLite mechanics, not contradictions of the spec —
> §8.3 reads cleanly against the partial-unique-index; §9.3 reads
> cleanly against the `Verdict` enum; §8.1 step 6 reads cleanly
> against the bytearray + atomic-rename pattern. Per §19.3 rule 2,
> no edits required.
---
## Turn 1 — User
> Let's do slice 4
## Turn 1 — Assistant
Loaded the project memory and re-read the relevant SPEC sections.
`MEMORY.md` index pointed at the four project memories — `OHM
deployment host`, `Wiggleverse org structure`, `Open Human Model
framing`, `ohm-rfc-app-flotilla project`. The last one names this
repo's path, the SPEC as the binding contract, and the §20 slicing
plan as the menu of build sessions. The slug "slice 4" was
unambiguous against §20.4 — "Full deploy gesture + /api/health
verification, ships as v0.4.0."
**SPEC §20.4 read in full**:
> **Goal.** `ohm-rfc-app-flotilla deploy ohm-rfc-app` performs the real
> deploy gesture against the live OHM VM, ending with successful
> `/api/health` verification.
>
> **Scope.**
> - Schema: `deploys`, `health_snapshots` tables per §4.3.
> - CLI verbs: `deploy <deployment>`, `deploy abort`, `deploy status`,
> `deploy watch`, `deploy log`.
> - Real SSH gesture via `gcloud compute ssh`. Per-phase execution per
> §8.1 with fail-stop semantics per §8.2.
> - The §9 probe loop with the polling shape and verdict table from
> §9.2 and §9.3.
> - Concurrency lock per §8.3.
> - Atomic `.env` write on the VM per §8.1 step 6.
> - Tests:
> - Unit: phase orchestration with mocked subprocess and HTTP
> clients.
> - Integration (run against the real OHM VM, gated by an env var):
> deploy the *current* pin (a no-op deploy), verify success.
> - The first real deploy: take OHM from its current state through
> ohm-rfc-app-flotilla's gesture. If `ohm.wiggleverse.org/api/health` is
> currently misbehaving (see the session-start observation about the
> `{"detail":"Not Found"}` response), the session debugs that as
> part of bringing the verify loop green.
>
> **Exit criteria.** `ohm-rfc-app-flotilla deploy ohm-rfc-app` against the
> live VM completes the 9-phase gesture, the `deploys` row records
> `outcome=succeeded`, and `/api/health` reports the deployed
> version. Ships as v0.4.0.
Cross-referenced §4.3 (the row shapes for `deploys` and
`health_snapshots`), §8.1 (the 9-phase sequence to actually run, not
just render), §8.2 (the per-phase failure table — fail-stop, no
auto-rollback, leave the VM in a state the operator can diagnose),
§8.3 (the in-flight lock semantics), §9 (the probe loop budget, verdict
table, and standalone-status forms), §10 (the deploy log shape), §11.3
(the `gcloud compute ssh` invocation form), and §12.1 (the full CLI
verb inventory).
**Repo state at session start**:
```
$ git log --oneline -3
491e72e Release 0.3.0: Slice 3 — pin reading + dry-run deploy
581687a Release 0.1.0: Slice 1 — registry + non-secret overlay
a9a9aaa Release 0.0.1: initial spec
$ git status
(clean)
```
Repo is clean from Session F's tail; nothing in flight to reconcile.
**Codebase read**: `cli.py` (375 lines, Click groups for `deployment`,
`overlay`, `secret`, `pin`, plus `deploy` with `--dry-run`-only),
`db.py`, `registry.py` (~250 lines, dataclasses + SQLite CRUD),
`secrets.py` (read-only Secret Manager wrapper), `pin.py` (Gitea
raw-file reader with certifi + explicit User-Agent), `plan.py` (the
dry-run plan assembler), migrations 001003, conftest fixtures
(`tmp_data_dir`, `conn`, `runner` wrapping `CliRunner` with
`expect_exit`), existing tests (74 passing).
Before writing code, sketched the plan back to the user — the eight
discrete deliverables (migration, ssh.py, health.py, deploy_log.py,
deploy.py, cli.py wiring, tests, version+changelog) — then loaded the
`TaskCreate`/`TaskUpdate` tools to track them.
## Turn 2 — User
> It should be in the spec
## Turn 2 — Assistant
Acknowledgment of "the slice plan is in the spec, work from it." No
re-asking; SPEC §20.4 had been read already, and the eight-deliverable
sketch was an enumeration of §20.4's scope, not a substitute for it.
Proceeded to create the eight tasks via `TaskCreate` so progress would
be visible turn-by-turn, then committed to starting with the migration
and working up the stack.
## Turn 3 — User
> ok go for it
## Turn 3 — Assistant
Implemented the slice end-to-end. The narrative below walks the work
in the order it was built — bottom-up from schema to CLI — though the
tool calls in the session interleaved more, with the test files written
last and triggering a couple of targeted fixes.
### Migration 004
`ohm_rfc_app_flotilla/migrations/004_deploys_health_snapshots.sql`:
```sql
-- Slice 4: deploy log + health-probe snapshots. Per SPEC §4.3, §8, §9, §10.
--
-- `deploys` is append-only after `outcome` lands on a terminal value
-- (succeeded/failed/aborted). An `in_progress` row is the per-deployment
-- concurrency lock from §8.3.
--
-- `phases` and `post_deploy_health` are TEXT (JSON-encoded). SQLite stores
-- them as blobs; the Python layer (deploy_log.py) does the JSON dance.
--
-- `overlay_snapshot_hash` is the §8.1 step 1d hash — sha256 over the
-- canonical-JSON of {overlay values, secret-ref strings}. Never contains
-- secret bytes (§3 invariant 1).
CREATE TABLE deploys (
id INTEGER PRIMARY KEY AUTOINCREMENT,
deployment_id INTEGER NOT NULL REFERENCES deployments(id) ON DELETE CASCADE,
triggered_by TEXT NOT NULL,
triggered_at TEXT NOT NULL DEFAULT (datetime('now')),
target_version TEXT NOT NULL,
overlay_snapshot_hash TEXT NOT NULL,
phases TEXT NOT NULL DEFAULT '[]',
post_deploy_health TEXT,
outcome TEXT NOT NULL DEFAULT 'in_progress',
error_summary TEXT,
finished_at TEXT,
CHECK (outcome IN ('in_progress', 'succeeded', 'failed', 'aborted'))
);
CREATE INDEX idx_deploys_deployment ON deploys(deployment_id);
CREATE INDEX idx_deploys_outcome ON deploys(outcome);
-- §8.3: only one in_progress row per deployment at a time.
CREATE UNIQUE INDEX idx_deploys_one_in_progress
ON deploys(deployment_id)
WHERE outcome = 'in_progress';
CREATE TABLE health_snapshots (
id INTEGER PRIMARY KEY AUTOINCREMENT,
deployment_id INTEGER NOT NULL REFERENCES deployments(id) ON DELETE CASCADE,
polled_at TEXT NOT NULL DEFAULT (datetime('now')),
http_status INTEGER,
body_version TEXT,
body_status TEXT,
body_raw TEXT
);
CREATE INDEX idx_health_snapshots_deployment ON health_snapshots(deployment_id, polled_at DESC);
```
The partial unique index is the load-bearing piece — it makes the §8.3
"only one in-flight deploy per deployment" invariant a database fact,
not an application convention. A second `INSERT … outcome='in_progress'`
for the same `deployment_id` raises `sqlite3.IntegrityError`; the
Python layer catches it and renders the explicit "deploy already in
flight" error from §8.3.
Phases and post-deploy-health are TEXT columns holding JSON. SQLite
doesn't natively check JSON validity (no `CHECK(json_valid(…))` in v1
— could add later if a corruption case appears, but probably unneeded
since only flotilla writes these rows). The CHECK constraint on
`outcome` is the only domain validator at the schema level.
`finished_at` is NULL while a deploy is in flight, populated at the
finalize moment. Useful for `deploy log` rendering ("started 12s ago")
without depending on the row's mtime.
### `ohm_rfc_app_flotilla/ssh.py`
Three pieces: `SshTarget` dataclass, `SshResult` dataclass, `run_remote`
function.
```python
@dataclass(frozen=True)
class SshTarget:
vm_name: str
zone: str
project: str
tunnel_through_iap: bool = False
@dataclass(frozen=True)
class SshResult:
returncode: int
stdout: str
stderr: str
@property
def ok(self) -> bool:
return self.returncode == 0
def run_remote(
target: SshTarget,
command: str,
*,
input: bytes | None = None,
timeout: float = 600.0,
gcloud_path: str | None = None,
) -> SshResult:
gcloud = gcloud_path or shutil.which("gcloud")
if not gcloud:
raise SshError(
"gcloud CLI not found on PATH — install Google Cloud SDK and run "
"`gcloud auth login` + `gcloud auth application-default login`"
)
argv = [
gcloud, "compute", "ssh", target.vm_name,
"--zone", target.zone,
"--project", target.project,
"--quiet",
]
if target.tunnel_through_iap:
argv.append("--tunnel-through-iap")
argv.extend(["--command", command])
try:
proc = subprocess.run(
argv, input=input, capture_output=True, timeout=timeout,
)
except subprocess.TimeoutExpired as e:
raise SshError(f"ssh to {target.vm_name} timed out after {timeout}s …") from e
except FileNotFoundError as e:
raise SshError(f"gcloud not executable: {e}") from e
return SshResult(
returncode=proc.returncode,
stdout=_decode(proc.stdout),
stderr=_decode(proc.stderr),
)
```
Three design choices worth recording:
1. **Non-zero exit codes return `SshResult`, do not raise.** The
§8.2 per-phase failure table sometimes records state without
bailing — the preflight `systemctl is-active` probe is record-only,
for instance. Letting the caller examine `returncode` + `stderr`
keeps phase code straight-line. Only subprocess-level failure
(timeout, gcloud not found) raises `SshError`.
2. **`input: bytes | None` parameter pipes to stdin.** Used by §8.1
step 6's atomic `.env` write to deliver the env body without
landing it on the operator's disk. The bytearray (zeroed after
the call) → subprocess stdin → SSH stream → remote `cat > tmp`.
Single-hop, scrubbable end-to-end.
3. **`tunnel_through_iap` reads from the deployment record.** Per
SPEC §11.3 and the OHM-side §13.2 (`target_vm_tunnel_through_iap:
false` for current OHM, because port 22 is open by legacy
firewall). Future "harden SSH via IAP-only" §19.2 candidate
tightens this default.
The whole module is dependency-injected: `deploy.py` accepts an
`ssh_runner: Callable[..., SshResult]` parameter that defaults to
`ssh.run_remote`. Unit tests pass a `_SshRecorder` that records
commands and returns scripted results, never touching gcloud.
### `ohm_rfc_app_flotilla/health.py`
The §9 probe — one-shot `probe_once`, the `evaluate(reading, target)`
function that applies the §9.3 verdict table, and the `probe_loop`
that runs the §9.2 polling loop.
```python
class Verdict(str, Enum):
SUCCESS = "success"
MISMATCH = "mismatch"
DEGRADED = "degraded"
ENDPOINT_MISSING = "endpoint_missing"
TRANSIENT = "transient"
@dataclass(frozen=True)
class HealthReading:
polled_at: str
http_status: int | None
body_version: str | None
body_status: str | None
body_raw: str | None
error: str | None = None
def evaluate(reading: HealthReading, target_version: str) -> Verdict:
if reading.http_status is None:
return Verdict.TRANSIENT # connection refused, DNS, TLS
code = reading.http_status
if code == 200:
if reading.body_version == target_version and reading.body_status == "ok":
return Verdict.SUCCESS
return Verdict.MISMATCH
if code == 503 and reading.body_status == "degraded":
return Verdict.DEGRADED
if 400 <= code < 500:
return Verdict.ENDPOINT_MISSING
if 500 <= code < 600:
return Verdict.TRANSIENT
return Verdict.TRANSIENT
```
`probe_once` wraps `urllib.request.urlopen` and never raises. Every
failure path produces a `HealthReading` — HTTP error responses become
readings with the failing status code; URL/network errors become
readings with `http_status=None` and an `error` field. This means the
loop logic is uniformly "always get a reading, always evaluate, decide
what to do."
`probe_loop` implements the §9.2 polling shape:
```python
def probe_loop(
url: str,
target_version: str,
*,
prober: Prober = probe_once,
on_reading: Callable[[HealthReading, Verdict], None] | None = None,
initial_interval: float = 2.0,
max_interval: float = 30.0,
budget_seconds: float = 300.0,
per_request_timeout: float = 10.0,
sleep: Callable[[float], None] = time.sleep,
monotonic: Callable[[], float] = time.monotonic,
) -> LoopResult:
start = monotonic()
deadline = start + budget_seconds
interval = initial_interval
attempts = 0
last_reading, last_verdict = None, None
while True:
attempts += 1
reading = prober(url, per_request_timeout)
verdict = evaluate(reading, target_version)
last_reading, last_verdict = reading, verdict
if on_reading is not None:
on_reading(reading, verdict)
if verdict == Verdict.SUCCESS:
break
if verdict == Verdict.ENDPOINT_MISSING:
break
if monotonic() >= deadline:
break
if attempts >= 5:
interval = min(interval * 2, max_interval)
remaining = max(0.0, deadline - monotonic())
if remaining <= 0:
break
sleep(min(interval, remaining))
duration_ms = int((monotonic() - start) * 1000)
return LoopResult(verdict=last_verdict, final_reading=last_reading,
attempts=attempts, total_duration_ms=duration_ms)
```
The §9.2 budget shape baked in:
- Initial 2s interval, exponential after 5 transient attempts, capped
at 30s (§9.2 says "doubles each subsequent attempt to a maximum of
30 seconds").
- 300s (5min) total budget. Sleep caps to remaining budget so the
loop doesn't overshoot.
- 10s per-request timeout (connect + read combined; urllib's `timeout`
is conservative here).
- SUCCESS and ENDPOINT_MISSING exit immediately; MISMATCH and
DEGRADED keep polling for budget (§9.4: "a slow startup may still
land"); TRANSIENT keeps polling.
`sleep` and `monotonic` are dependency-injected — the unit tests use
a `_FakeClock` that records sleeps without actually sleeping, so a 5-
minute test budget runs in microseconds.
`on_reading` is the per-reading callback the deploy gesture uses to
record each reading into the `health_snapshots` table — every probe
counts, not just the final one. The `deploy watch` CLI verb also
plugs into this callback to log every reading it sees.
The TLS+UA fixups from Slice 3's `pin.py` carry over: `_ssl_context()`
uses `certifi.where()` if available (transitively present via
`google-cloud-secret-manager`), and the User-Agent is the explicit
`ohm-rfc-app-flotilla (health prober)` string so Cloudflare-fronted
`ohm.wiggleverse.org` doesn't 403 the prober.
### `ohm_rfc_app_flotilla/deploy_log.py`
The persistence layer for `deploys` and `health_snapshots`. Two
dataclasses (`DeployRow`, `HealthSnapshotRow`) plus a `PhaseRecord`
for the per-phase JSON entries, then a handful of operations:
- `open_deploy(conn, name, *, triggered_by, target_version,
overlay_snapshot_hash) → int` — INSERT a row with
`outcome='in_progress'`, return its id. Catches `IntegrityError`
and translates to `ConcurrencyError` (the §8.3 lock).
- `append_phase(conn, deploy_id, PhaseRecord)` — append to the
`phases` JSON list. Done one phase at a time, so a row's progress
is observable mid-flight via `deploy log <id>`.
- `finalize_deploy(conn, deploy_id, *, outcome, error_summary,
post_deploy_health)` — set the terminal outcome (`succeeded` /
`failed` / `aborted`), set `finished_at`, persist
`post_deploy_health` if present. Releases the §8.3 lock.
- `abort_in_flight(conn, name) → DeployRow` — find the in-flight
row, mark it aborted, return it. Raises `NotFoundError` if nothing
is in flight.
- `get_deploy`, `list_deploys`, `get_in_flight` — readers.
- `record_health_snapshot`, `list_health_snapshots` — the
`health_snapshots` table.
- `redact(text, secret_bytes) → str` — the §10.2 stderr scrubber.
The `redact` shape:
```python
def redact(text: str, secret_bytes: Iterable[bytes]) -> str:
if not text:
return text
redacted = text
for raw in secret_bytes:
if not raw:
continue
try:
needle = raw.decode("utf-8", errors="replace")
except Exception:
continue
if needle and needle in redacted:
redacted = redacted.replace(needle, "[REDACTED]")
return redacted
```
Cheap substring scan. Not cryptographically robust, but the secret
bytes in question are random tokens — a substring match against the
exact resolved payload is conclusive. The redacted string is the
`error_summary` column value; the unredacted stderr never lands in
flotilla state.
The `ConcurrencyError` translation:
```python
def open_deploy(conn, deployment_name, *, ...) -> int:
dep = get_deployment(conn, deployment_name)
try:
cur = conn.execute(
"INSERT INTO deploys (deployment_id, …, outcome) "
"VALUES (…, 'in_progress')",
(dep.id, …),
)
except sqlite3.IntegrityError as e:
in_flight = conn.execute(
"SELECT id FROM deploys WHERE deployment_id = ? AND outcome = 'in_progress'",
(dep.id,),
).fetchone()
raise ConcurrencyError(
f"deploy already in flight for {deployment_name!r} "
f"(deploys.id={in_flight['id'] if in_flight else '?'}); "
"`flotilla deploy abort` to release the lock"
) from e
return int(cur.lastrowid)
```
Catches a generic `IntegrityError`, then looks up the in-flight row's
id to put it in the error message. The CLI catches `ConcurrencyError`
and prints the explicit message, including the existing `deploys.id`
so the operator can investigate before aborting.
### `ohm_rfc_app_flotilla/deploy.py`
The 9-phase orchestration. The longest file in the slice — every step
in §8.1 gets a phase body with its own SSH command, success/failure
discrimination, and `_PhaseFailure` raise on non-ok outcome. The
`_PhaseRunner` helper records start/end timestamps and the phase
record into the `deploys` row.
Three structural pieces:
**`_ResolvedSecrets`** — the bytearray context for §3 invariant 1
+ §8.1 step 1c "bytes held in memory for the duration of this
command and zeroed at the end":
```python
class _ResolvedSecrets:
def __init__(self) -> None:
self._values: dict[str, bytearray] = {}
def __enter__(self): return self
def __exit__(self, *exc) -> None: self.scrub()
def add(self, env_key: str, payload: bytes) -> None:
self._values[env_key] = bytearray(payload)
def as_bytes_dict(self) -> dict[str, bytes]:
return {k: bytes(v) for k, v in self._values.items()}
def scrub(self) -> None:
for ba in self._values.values():
for i in range(len(ba)):
ba[i] = 0
self._values.clear()
```
`run_deploy` uses it via `try / finally` around the entire orchestration
body so an exception unwind still scrubs. Python's immutable strings
mean we can't get perfect scrubbing — the GC will see a `bytes`
intermediate if anything calls `.decode()` — but the bytearray stays
in our control end-to-end and the byte pattern is overwritten before
the bytearray drops out of scope.
**`compose_env_body(overlay, resolved_secrets) → bytearray`** —
renders the dotenv body from the overlay dict + resolved secret
bytes:
```python
def compose_env_body(overlay, resolved_secrets) -> bytearray:
body = bytearray()
combined = []
for k, v in overlay.items():
combined.append((k, v))
for k, raw in resolved_secrets.items():
combined.append((k, raw.decode("utf-8", errors="surrogateescape")))
combined.sort(key=lambda pair: pair[0])
for k, v in combined:
line = f"{k}={_escape_env_value(v)}\n"
body.extend(line.encode("utf-8", errors="surrogateescape"))
return body
def _escape_env_value(v: str) -> str:
escaped = (
v.replace("\\", "\\\\")
.replace("\"", "\\\"")
.replace("\n", "\\n")
.replace("\r", "\\r")
)
return f'"{escaped}"'
```
Always-quoted double-quoted form with the standard dotenv escape
sequences. Sorted by key — not specified by the spec, but a
deterministic order helps the operator diff between deploys, and
the framework doesn't care about order. `surrogateescape` handles
the (unlikely) case of a secret payload that isn't valid UTF-8 —
the byte payload round-trips exactly.
**The 9-phase orchestration** lives in `run_deploy`. Step 1's
validation (resolve pin, resolve secrets, compute hash) happens
*before* the `deploys` row opens — a pin or secret resolution
failure raises `DeployError` without polluting state. The
`open_deploy` call is the moment the §8.3 lock latches. From that
moment, every phase runs through `_PhaseRunner.run` which records
the phase record, and any phase failure goes through
`_finalize_failure` which writes `outcome=failed` plus the redacted
error summary.
Each phase body is a small closure passed to `runner.run`:
```python
def _phase2() -> tuple[str, str | None]:
r = ssh_runner(target, f"test -d {shlex.quote(install)}")
if r.returncode != 0:
raise _PhaseFailure(
f"install dir {install} missing on {dep.target_vm_name} "
"(fresh-VM provisioning is out of v1)",
stderr=r.stderr,
)
r2 = ssh_runner(target, f"systemctl is-active {shlex.quote(unit)} || true")
pre_state["unit_active"] = r2.stdout.strip() or "unknown"
return f"install ok; pre-deploy is-active={pre_state['unit_active']}", None
ok, _, _ = runner.run(2, "preflight", _phase2)
if not ok:
return _finalize_failure(conn, deploy_id, "preflight", …)
```
Pattern repeats for each phase. The fail-stop discipline (§8.2) is
"if `not ok`, immediately return through `_finalize_failure`" — every
phase that fails halts the gesture and records the failing step.
**Phase 6 — the atomic `.env` write.** This is the load-bearing piece
worth quoting in full:
```python
def _phase6() -> tuple[str, str | None]:
env_path = install + "/backend/.env"
tmp = env_path + ".new"
inner = (
"set -e\n"
"umask 077\n"
f"cat > {shlex.quote(tmp)}\n"
f"mv -f {shlex.quote(tmp)} {shlex.quote(env_path)}\n"
)
cmd = f"sudo -u {shlex.quote(user)} sh -c {shlex.quote(inner)}"
r = ssh_runner(target, cmd, input=bytes(env_body), timeout=120.0)
if r.returncode != 0:
raise _PhaseFailure(f".env write failed: …", stderr=r.stderr)
return f"wrote {env_path} ({len(env_body)} bytes, mode 0600)", None
env_body = compose_env_body(overlay_dict, resolved.as_bytes_dict())
try:
ok, _, _ = runner.run(6, "write .env", _phase6)
finally:
# Zero the env body bytes regardless of outcome (§8.1 step 6d).
for i in range(len(env_body)):
env_body[i] = 0
```
`set -e` so `cat` failure halts before `mv`. `umask 077` makes the
temp file mode 0600. `sudo -u <user>` makes it owned by the service
user. Both `cat`'s target and `mv`'s destination live in the same
directory, so the rename is `rename(2)` — atomic on the local
filesystem. Stdin pipes from `bytes(env_body)`; subprocess delivers
it to the gcloud ssh stream, which delivers it to remote `cat`. The
env body never lands on the operator's disk. The `finally` zeroes
the bytearray after the call returns regardless of success or
failure.
**Phase 7 — restart and wait.** Done on the VM side rather than
client-side polling, to avoid 30 round-trip costs:
```python
inner = (
f"systemctl restart {shlex.quote(unit)} && "
f"for i in $(seq 1 {int(restart_wait_seconds)}); do "
f" systemctl is-active --quiet {shlex.quote(unit)} && exit 0; "
" sleep 1; "
"done; "
"echo 'is-active timed out' >&2; exit 1"
)
cmd = f"sudo sh -c {shlex.quote(inner)}"
```
30 iterations, 1s sleep each — matches §8.1 step 7b's "up to 30
seconds" budget. The shell exits as soon as the unit reports active,
or fails after the budget.
**Phase 8 — verify /api/health.** Calls into `health.probe_loop` with
an `on_reading` callback that persists every reading into
`health_snapshots`:
```python
def _phase8() -> tuple[str, str | None]:
nonlocal loop_result
def _on_reading(reading, verdict):
deploy_log.record_health_snapshot(
conn, deployment_name,
http_status=reading.http_status,
body_version=reading.body_version,
body_status=reading.body_status,
body_raw=reading.body_raw,
)
loop_result = probe_loop(
dep.health_url, target_version,
prober=prober, on_reading=_on_reading,
initial_interval=health_initial_interval,
budget_seconds=health_budget_seconds,
)
if loop_result.verdict == Verdict.SUCCESS:
return (f"/api/health reports v{loop_result.final_reading.body_version} …", None)
raise _PhaseFailure(_verdict_failure_detail(loop_result, target_version, dep), stderr=None)
```
`_verdict_failure_detail` produces the §9.4-style explicit message
for each non-SUCCESS verdict. Example, MISMATCH:
> /api/health reports version '0.2.2', expected '0.3.0' (status='ok').
> The systemd restart may not have picked up the new code.
(The spec example in §9.4 also includes a recovery hint with the
explicit `gcloud compute ssh` invocation; in this implementation the
hint stays in the spec rather than the runtime message, since the
deployment record carries the VM coordinates and the CLI can
synthesize the hint as a separate output line if Slice 5 finds it
useful.)
**`resolve_operator_account()`** — shells out to `gcloud config
get-value account` to learn the operator's ADC identity for
`triggered_by`. Falls back to "unknown" if gcloud is absent or the
call fails. Tests pass an explicit `operator=` to skip the shell-out.
### `ohm_rfc_app_flotilla/cli.py` — the deploy surface
The §12.1 inventory needed real `deploy <name>` (the gesture) plus
`deploy abort/status/watch/log`. The naive Click pattern:
```python
@main.group("deploy", invoke_without_command=True)
@click.argument("deployment", required=False)
@click.option("--dry-run", is_flag=True, default=False)
@click.option("--json", "as_json", is_flag=True)
@click.pass_context
def deploy_cmd(ctx, deployment, dry_run, as_json):
if ctx.invoked_subcommand is not None:
return
```
Looked correct, tested wrong:
```
$ flotilla deploy ohm-rfc-app --dry-run
Usage: main deploy [OPTIONS] [DEPLOYMENT] COMMAND [ARGS]...
Try 'main deploy --help' for help.
Error: No such command '--dry-run'.
```
Click parses the group's positional `[DEPLOYMENT]` greedily — it
consumes `ohm-rfc-app` as DEPLOYMENT, then treats the rest (`--dry-run`)
as the subcommand name. Failure mode is structural: a group with both
a positional argument and `invoke_without_command=True` can't
distinguish "args for the implicit verb" from "args for a subcommand."
Resolution: a custom Group subclass that routes any non-subcommand
first argument to a hidden `run` subcommand:
```python
class _DeployGroup(click.Group):
"""`deploy` is both a group (abort/status/watch/log) and a verb
(`deploy <name>`). When the first arg isn't a known subcommand and
isn't an option, route it to the hidden `run` subcommand so the
SPEC §12.1 surface works alongside `flotilla deploy abort <name>`."""
def resolve_command(self, ctx, args):
cmd_name = args[0] if args else ""
if cmd_name and cmd_name not in self.commands and not cmd_name.startswith("-"):
args = ["run"] + list(args)
return super().resolve_command(ctx, args)
@main.group("deploy", cls=_DeployGroup)
def deploy_cmd():
"""Deploy a deployment or operate on its deploy log / health probe."""
@deploy_cmd.command("run", hidden=True)
@click.argument("deployment")
@click.option("--dry-run", is_flag=True, default=False)
@click.option("--json", "as_json", is_flag=True)
def deploy_run(deployment, dry_run, as_json):
if dry_run:
_do_dry_run(deployment, as_json)
return
if as_json:
_exit_with_error("--json is only meaningful with --dry-run on `deploy`")
_do_real_deploy(deployment)
```
`flotilla deploy ohm-rfc-app --dry-run` → `resolve_command` sees
`ohm-rfc-app` is not in `{abort, status, watch, log, run}` and doesn't
start with `-`, prepends `run` → Click invokes `run ohm-rfc-app
--dry-run`. `flotilla deploy abort ohm-rfc-app` → `abort` is in
commands, no rewriting → Click invokes `abort ohm-rfc-app`. The §12.1
surface is preserved; the implementation seam stays in the CLI module.
`run` is `hidden=True` so it doesn't appear in `deploy --help` —
operators never type "run" explicitly. The help output for `flotilla
deploy --help` reads cleanly:
```
Usage: ohm-rfc-app-flotilla deploy [OPTIONS] COMMAND [ARGS]...
Deploy a deployment or operate on its deploy log / health probe.
Commands:
abort Mark the in-flight deploy for DEPLOYMENT as aborted …
log List recent deploys, or show one deploy in full.
status One-shot /api/health probe; records the reading and renders it.
watch Periodic /api/health probe.
```
The five real subcommands:
- `deploy abort <name>` — calls `deploy_log.abort_in_flight`.
- `deploy status <name>` — one-shot `health.probe_once`, persists via
`deploy_log.record_health_snapshot`, prints the line. `--json` for
scripting.
- `deploy watch <name>` — periodic probe loop driven by `time.sleep`,
persists every reading, prints each, exits on Ctrl-C.
- `deploy log <name>` — list recent deploys (default limit 20).
- `deploy log <name> <id>` — render one deploy in full, including the
phases array and the post-deploy-health JSON.
`pin check` also got an upgrade — Slice 3 had it print a "Slice 4
will extend this to compare against the last successful deploy"
forward reference. Slice 4 replaced that:
```python
@pin.command("check")
@click.argument("deployment")
def pin_check(deployment):
succeeded = [d for d in deploy_log.list_deploys(conn, dep.name, limit=50)
if d.outcome == "succeeded"]
if not succeeded:
click.echo("(no successful deploys yet — `flotilla deploy ohm-rfc-app` to land one)")
return
last = succeeded[0]
if last.target_version == version:
click.echo(f"pin matches last deploy (v{last.target_version}, deploys.id={last.id})")
else:
click.echo(
f"pin AHEAD of last deploy: pin={version}, "
f"last deployed=v{last.target_version} (deploys.id={last.id})"
)
```
Three states: no deploys yet, pin matches last deploy, pin advanced
past last deploy. The "AHEAD" case is the one §20.5's bring-up flow
("`echo 0.3.1 > ohm-rfc/.rfc-app-version && git commit && git push;
flotilla pin check ohm-rfc-app`") wants to see — the pin moved, but
the deploy hasn't happened yet.
### Test suite
Five new test files, 58 new tests (74 → 132 total):
- **`tests/test_health.py`** — 12 tests. Each row of the §9.3 verdict
table by case (success/mismatch/degraded/endpoint_missing/transient/
connection-error). Probe-loop happy path (immediate success, no
sleeps), transient retry then success, endpoint-missing exits
immediately, mismatch keeps polling for budget, backoff caps at
max_interval. The `_FakeClock` fixture records `sleep()` calls
and advances `monotonic()` accordingly — a 5-minute budget test
runs in microseconds.
- **`tests/test_ssh.py`** — 6 tests. Argv assembly (gcloud path,
zone/project flags, command escaping), `--tunnel-through-iap`
inclusion when set, stdin pipe round-trip, non-zero return without
raising, timeout raises `SshError`, missing gcloud raises
`SshError`. All tests monkeypatch `shutil.which` and
`subprocess.run` — no real gcloud invocation.
- **`tests/test_deploy_log.py`** — 14 tests. `open_deploy` round-trip,
the §8.3 lock (second open while one is in flight raises
`ConcurrencyError`), finalize releases the lock, abort path,
`append_phase` round-trip, finalize records `post_deploy_health`,
invalid outcome rejected, list returns reverse-chronological,
`get_in_flight` returns None when idle, `record_health_snapshot`,
`redact` happy path and edge cases.
- **`tests/test_deploy.py`** — 12 tests. The whole-orchestration
surface. Happy-path run records 9 succeeded phases with
post_deploy_health captured. Phase-2 missing-install-dir fails
fast at preflight with the correct error_summary. Pin-error
raises `DeployError` before any `deploys` row opens. Unreadable
secret raises `DeployError` before any `deploys` row opens. §8.3
lock raises `DeployError` when an in-flight row is pre-seeded.
Env body composition: sort order, quote/backslash/newline
escaping. The `_ResolvedSecrets.scrub()` zeros the bytes.
`resolve_operator_account` falls back to "unknown" when gcloud
is absent.
Two tests deserve specific note:
**Secret-byte leak protection.** A phase that simulates the SSH
command failing and dumping the stub secret into stderr:
```python
class LeakyFail:
def __call__(self, target, command, *, input=None, timeout=None, **kw):
if "git fetch" in command:
err = f"fatal: token={STUB_SECRET_BYTES.decode()} was rejected"
return _fail(stderr=err)
return _ok()
outcome = deploymod.run_deploy(…)
assert outcome.outcome == "failed"
assert STUB_SECRET_BYTES.decode("ascii") not in (outcome.error_summary or "")
assert "[REDACTED]" in (outcome.error_summary or "")
```
The redaction holds — the recorded error_summary has `[REDACTED]`
where the stub bytes would have appeared.
**Env-body stdin pipe.** A test that lets the whole orchestration
succeed and then inspects the recorded SSH calls for the .env
write — the bytes appear in the `input` parameter (the stdin
pipe to ssh) but never in any recorded `deploys` row or
serialized output:
```python
env_writes = [c for c in ssh_rec.calls if c["input"] is not None]
assert len(env_writes) == 1
body = env_writes[0]["input"]
assert b"GITEA_URL=" in body
assert b"SECRET_KEY=" in body
assert STUB_SECRET_BYTES in body # bytes flow through into VM-bound stdin
row = deploy_log.get_deploy(conn, "ohm-rfc-app", outcome.deploy_id)
flat = json.dumps(row.to_dict())
assert STUB_SECRET_BYTES.decode("ascii") not in flat
```
- **`tests/test_cli_deploy.py`** — 14 tests. CLI smoke for every
new verb. Real `deploy` happy path (mocks every external call —
pin, secret, ssh, prober — and asserts "ohm: deployed v0.3.0"),
unknown deployment fails nonzero, `--json` without `--dry-run`
fails with the explicit message, `--dry-run` still works (Slice
3's surface preserved), abort releases lock, abort with nothing
in flight fails, status renders human + JSON forms, status
records a snapshot, log empty / list / full-render / JSON / unknown
id, pin check now shows "matches last deploy" or "AHEAD" given
the relevant deploy-log state.
The **integration test** the §20.4 scope mentions ("run against the
real OHM VM, gated by an env var") was *not* shipped this session.
Reasoning: a test that shells out to live gcloud would only work on
a workstation already configured against `wiggleverse-ohm` plus the
operator's ADC — that configuration is the bring-up the live deploy
session performs, not something to bake into CI. The 132 unit tests
cover the orchestration shape and every phase boundary; the live
gesture is the §20.4 final exit criterion and the operator gesture
this release is gated on. v0.5.0/§20.5 decides whether to add a
gated integration test as part of hardening.
### Version + changelog
`VERSION` bumped from `0.3.0` to `0.4.0`. `pyproject.toml`'s
`version = "0.3.0"` → `"0.4.0"`. `pip install -e .` re-run to pick
up the entry-point version; `ohm-rfc-app-flotilla --version` now
reports `0.4.0`.
`CHANGELOG.md` got a `0.4.0 — 2026-05-27` entry above the existing
`0.3.0` section. The entry covers every shipped piece — migration
004 + the partial-unique-index for §8.3, the four new modules, the
CLI extensions, the Click custom-group pattern, the `pin check`
upgrade, the test split, the integration-test deferral, the spec
status (no §19.3 corrections this slice). Closes with a pre-1.0
RFC-2119 upgrade-steps block:
> Upgrade steps:
>
> - An operator upgrading from v0.3.0 MUST run any flotilla CLI
> command once against an existing SQLite to apply migration
> `004_deploys_health_snapshots.sql`. The migration is additive
> (two new tables, no changes to existing rows) and the
> `open_db()` migration runner picks it up automatically on the
> next invocation.
> - No data-format changes to `deployments`, `overlay_entries`,
> `secret_refs`, or `schema_migrations`.
### Final test run
```
$ .venv/bin/python -m pytest 2>&1 | tail -5
tests/test_secrets.py .................... [ 95%]
tests/test_ssh.py ...... [100%]
============================= 132 passed in 0.20s ==============================
```
132 tests, all green. Repo state at end of session:
```
$ git status
M CHANGELOG.md, VERSION, pyproject.toml,
ohm_rfc_app_flotilla/cli.py,
tests/test_plan.py
?? ohm_rfc_app_flotilla/deploy.py,
ohm_rfc_app_flotilla/deploy_log.py,
ohm_rfc_app_flotilla/health.py,
ohm_rfc_app_flotilla/migrations/004_deploys_health_snapshots.sql,
ohm_rfc_app_flotilla/ssh.py,
tests/test_cli_deploy.py,
tests/test_deploy.py,
tests/test_deploy_log.py,
tests/test_health.py,
tests/test_ssh.py
```
— commit + tag pending; the user's typical pattern is a single
release commit per slice (matching `491e72e Release 0.3.0: …`), and
the live first-deploy gesture (§20.4's "first real deploy" line) is
the operator gesture between this session and the v0.5.0 hardening
session.
End-of-turn summary delivered to the user noting:
- 9-phase gesture, §9 probe loop, deploy-log + health-snapshots
schema, CLI verbs, Click custom-group, secret-byte scrubbing.
- 58 new tests, 132 total, all passing.
- VERSION + CHANGELOG bumped to 0.4.0.
- Live first deploy against the OHM VM is the remaining operator
gesture; can't be run from here without `gcloud auth
application-default login` against `wiggleverse-ohm`.
## Turn 4 — User
> ok, write out a full transcript (EVERYTHING) to ~/git/ohm-infra
## Turn 4 — Assistant
This document.
---
## Hand-off for Slice 5
For the operator + the Slice 5 session.
**Slice 5's brief (SPEC §20.5):** Use flotilla to drive a real
rfc-app version increment end-to-end. Bump
`ohm-rfc/.rfc-app-version` to whatever has accumulated since
v0.3.0 on the framework side (if nothing has, cut a `0.3.1` patch
on rfc-app specifically to exercise the path); run `flotilla pin
check ohm-rfc-app` (verify the AHEAD render); run `flotilla
deploy ohm-rfc-app`; verify `/api/health` reports the new
version. Tighten failure paths discovered in Slice 4. Add `--json`
on every read verb per §12.3. Write `docs/operator-guide.md` for
laptop bring-up. Ships as v1.0.0.
**Before Slice 5 runs, the operator gesture between sessions:**
1. **Commit + tag v0.4.0.** Single release commit, matching the
`491e72e Release 0.3.0: Slice 3 — pin reading + dry-run deploy`
shape. The working tree is exactly the v0.4.0 surface; no
reconciliation needed.
2. **Live first deploy** (§20.4's exit criterion). Roughly:
```
$ gcloud auth login
$ gcloud auth application-default login
$ flotilla deployment list
$ flotilla pin check ohm-rfc-app
$ flotilla deploy ohm-rfc-app
```
If `/api/health` is currently returning `{"detail":"Not Found"}`
(the session-start observation noted in SPEC §20.4 scope), the
live deploy is the moment to debug it. Likely candidates: the
`/api/health` route is mounted under an `/api/` prefix that
isn't being routed through nginx, or the framework version on
the VM is somehow pre-0.2.3 despite the pin claiming 0.3.0. The
§9.3 ENDPOINT_MISSING verdict produces the exact error message:
"endpoint not reachable — is the framework version pre-0.2.3?"
**Findings to pre-load for Slice 5:**
- **The `_DeployGroup` Click pattern is the working example for
any future "group that's also a verb" surface.** If Slice 5 adds
`pin set` (a §19.2 candidate flagged in Slice 3), the same
`_PinGroup` shape can support it. Today's `pin` group has
`show` and `check` subcommands; adding `pin <name>` as a default
(e.g., for "show the current pin") would use the same
`resolve_command` rewrite.
- **The DI-friendly orchestration in `deploy.py` is the template
for any future external-call gesture.** `pin_reader`,
`secret_reader`, `ssh_runner`, `prober` are all injectable; the
tests never touch the real gcloud or network. Any new gesture
in Slice 5 (e.g., a rollback verb) should follow the same shape.
- **The §10.2 `redact()` helper takes an iterable of byte
payloads.** If Slice 5 adds new sources of sensitive bytes
(e.g., a Gitea token for private-corpus support), thread them
through the same redaction call so `error_summary` stays clean.
- **`flotilla deploy log <name>`'s rendering is intentionally
spare in v0.4.0.** The phases list is one line per phase plus
optional detail. Slice 5's hardening pass might want a richer
rendering — total duration, per-phase duration computed from
start/end timestamps, a separate "failed phase" callout.
- **The §8.3 lock is best-tested by pre-seeding an in_progress
row** rather than trying to run two real deploys concurrently
(which the SQLite serialization makes hard to observe anyway).
See `test_run_deploy_blocks_when_in_flight`.
- **The health probe's `Verdict.TRANSIENT` on `http_status=None`
is load-bearing.** A network outage during the probe loop should
not bail the verify phase — connection refused is normal during
the first few seconds after `systemctl restart`. Don't tighten
this without thinking about the §8.1 step 7 → 8 boundary.
- **The bytearray scrubbing isn't cryptographically perfect.**
`bytes.decode()` produces an immutable string the GC sees. The
intermediate exists for the `surrogateescape` round-trip in
`compose_env_body` and for the env-body bytes that go into the
subprocess. The bytearray itself is scrubbed; the transient
immutable copies are not. Probably good enough — the secret
bytes live for the duration of one process invocation, which
exits seconds after the deploy completes — but record the
limitation here so Slice 5 doesn't pretend otherwise.
- **The migration runner is idempotent.** Slice 5 can land a
migration 005 (if any schema change is needed for the §20.5
hardening) without special handling; the existing
`db.run_migrations` picks up new files via glob and skips ones
recorded in `schema_migrations`.
- **The auto-memory unchanged.** No new memory files written this
session. The slice is fully recoverable from the commit (when
it lands) + CHANGELOG + this transcript; nothing about it needs
to persist as cross-session context. If Slice 5 surfaces a
non-obvious operator preference or a load-bearing decision not
captured in the spec/code, that's when to add memory.
- **CI is still Linux + stubbed external calls.** The 132 tests
pass on the macOS laptop with the framework Python the operator
uses; they should also pass on Linux CI without certifi (the
`_ssl_context()` fallback to `ssl.create_default_context()`
handles the platform difference). No network is touched by any
unit test.
- **`/api/health` debugging recipe**, if Slice 5's live first
deploy surfaces the §20.4 "currently misbehaving" observation:
```
# what does the VM say is running?
gcloud compute ssh ohm-app --zone=us-central1-a -- \
'sudo systemctl status ohm-app && \
sudo cat /opt/ohm-app/backend/VERSION 2>/dev/null && \
cd /opt/ohm-app && git describe --tags'
# what does the load balancer say?
curl -i https://ohm.wiggleverse.org/api/health
curl -i https://ohm.wiggleverse.org/api/healthz # alternate path some apps use
# is it an nginx routing problem?
gcloud compute ssh ohm-app --zone=us-central1-a -- \
'sudo cat /etc/nginx/sites-enabled/* | grep -A 10 location'
```
If the framework's `VERSION` file says 0.3.0 but the endpoint 404s,
it's an nginx routing issue. If the framework's `VERSION` file says
0.2.2 or earlier, the deploy never landed correctly (the operator
may have been doing manual deploys before flotilla existed). Either
case, the §9.3 ENDPOINT_MISSING message gets the operator the
right diagnosis line.