# Session 0028.0 — Transcript > Date: 2026-05-28 > Start: 2026-05-28T14-05 (PST implied) > End: 2026-05-28T14-32 (PST implied) > Subject: flotilla v1.2.0 — deploy robustness + record-maintenance verbs (PR, not merged) > Status: **FINALIZED.** --- ## Plan **Item picked:** flotilla **v1.2.0** — deploy-robustness + record-maintenance verbs (the §19.2 infra candidate surfaced by Session 0022). Pure-flotilla, touches no rfc-app code → zero collision with the in-flight rfc-app feature sessions 0026/0027. No OHM deploy, no version-slot contention, no secret bytes. Working on branch `session-0028/flotilla-v1.2.0`; shipping as a **PR** (operator merges + tags — per the "let's use branches and PRs" directive this session). Deliverables: 1. **Preflight SSH-failure disambiguation** (deploy.py phase 2). Today any non-zero `test -d ` is reported "install dir missing" — but a `gcloud compute ssh` connection/auth/IAP failure also returns non-zero, so a connectivity failure is misreported as a missing dir (deploy.py:355). Fix: a connectivity probe runs first; its failure reports the real cause. 2. **SshError → clean phase failure.** `_PhaseRunner.run` only catches `_PhaseFailure`; an `ssh.SshError` (e.g. a restart SSH timeout) propagates and leaves the `deploys` row stuck `in_progress` holding the §8.3 lock. Catch it, record a phase failure, finalize the row. 3. **SSE pre-drain restart** (phase 7). Replace `systemctl restart` with a flotilla-controlled stop (SIGTERM → bounded drain window → SIGKILL fallback for lingering SSE keepalive streams) → start. Bounds the window flotilla waits regardless of the unit's TimeoutStopSec. This is the v0.6.0 (Session 0009/"I") failure made un-reproducible from the flotilla side. A graceful app-cooperative drain endpoint in rfc-app stays a future §19.2 candidate. 4. **`flotilla deployment update`** verb (registry + CLI) — update mutable record fields without remove+add. Motivated by the #1 VM rename. 5. **`flotilla deploy reconcile`** verb — probe `/api/health` and reconcile the deploy log to reality: finalize a stuck `in_progress` row as `succeeded` when health confirms the target version (releases the lock); honestly report terminal rows health contradicts without falsifying the audit trail. Each lands with unit tests (no network — the deploy seams are DI'd). CHANGELOG 1.2.0 + VERSION/pyproject bump. Transcript finalized + published LAST. Parallel context at claim: sessions 0026 + 0027 also in-progress (launch prompts not captured in their placeholders). flotilla-only scope chosen specifically to stay disjoint from whatever rfc-app surfaces they touch. --- ## Friction log - **14:09 — shared-working-directory collision.** Branches alone don't isolate concurrent sessions: git tracks one working tree per checkout. Shortly after I branched in the canonical flotilla dir, another live session (0026 or 0027) wrote `ohm_rfc_app_flotilla/deploy.py` + `tests/test_deploy.py` in the **same** directory (a secret-redaction hardening, "finding L1"). Their files overlap exactly with mine (`_PhaseRunner`, phase details). Editing in place would corrupt both sessions' work. Resolution: restored the shared dir to `main` with their unstaged edits preserved, then created a **manual git worktree** at `…/ohm-rfc-app-flotilla-s0028` on my branch with its own fresh venv (`pip install -e .[test]`; the canonical `.venv`'s PEP 660 editable finder is pinned to the shared dir, so a shared venv would defeat isolation). Worktree baseline: 157 tests green on clean `c26f251`. **Handoff implication:** parallel flotilla sessions MUST each take a worktree; the operator's "branches and PRs" guidance needs the worktree step to actually isolate. Their redaction change and my v1.2.0 change both touch `deploy.py` → expect a merge conflict in `_PhaseRunner` when the second PR merges; resolvable (different concerns: they redact phase detail, I add an `SshError` branch + DI redactor coexist cleanly). ## Launch prompt You are an OHM roadmap driver session, running in parallel with others. Claim session ID first. STATE (end of 0022): OHM live at rfc-app v0.24.0 on the renamed VM `ohm-rfc-app`; deploys unblocked. Pick ONE disjoint roadmap item. Strong infra candidate: flotilla v1.2.0 — fix the preflight that misreports an SSH/auth failure as "install dir missing" (deploy.py:355), add a `deployment update` verb, add `flotilla deploy reconcile` + SSE pre-drain (§19.2 from this session). "Let's use branches and PR's so we don't need to worry about conflicting with other sessions." --- ## What shipped flotilla **v1.2.0**, on branch `session-0028/flotilla-v1.2.0`, pushed to `origin` (gitea). **Open as a PR (not yet merged):** https://git.wiggleverse.org/wiggleverse/ohm-rfc-app-flotilla/pulls/new/session-0028/flotilla-v1.2.0 Commit `51c7ea8`. No OHM deploy, no rfc-app release, no pin bump — this is operator-tool code. Operator opens + merges the PR (sole-merger rule) and tags `v1.2.0` after merge. Five changes, all flotilla-internal, each unit-tested (the deploy seams are dependency-injected so nothing touches gcloud or the network): 1. **Preflight SSH-failure disambiguation** (`deploy.py` phase 2). The misreport at `deploy.py:355`: a bare `test -d ` whose any-non-zero return was read as "install dir missing" — but `gcloud compute ssh` exits non-zero on a connection/auth/IAP failure too, so an unreachable VM was misdiagnosed as a missing dir. Fix: a trivial `true` connectivity probe runs first; its non-zero return is reported as the auth/connectivity failure it is. (Initially gold-plated with a stdout marker; dropped that for returncode-only after it would have forced every test stub — and every future one — to simulate the echo. Returncode-only fully solves the actual bug.) 2. **`_PhaseRunner` catches `ssh.SshError`** → records a clean `failed` phase instead of unwinding past `run_deploy` and stranding the row `in_progress` holding the §8.3 lock. This is the structural cause of the v0.6.0 / Session 0009 "aborted but actually fine" — a restart SSH timeout raised mid-phase. 3. **Phase-7 SSE pre-drain.** Replaced `systemctl restart` with a flotilla-controlled `systemctl kill -s TERM` → bounded `drain_wait_seconds` (default 20s) → `systemctl kill -s KILL` for lingering SSE keepalive streams → `systemctl start`. Bounded by `drain + restart_wait` regardless of the unit's `TimeoutStopSec`, so a held-open SSE stream can't blow past the SSH timeout. The phase detail records whether it drained via graceful SIGTERM or had to SIGKILL. 4. **`flotilla deployment update [--flag …]`** (registry + CLI). In-place edit of mutable record fields (VM name/zone/project/service user/install dir/systemd unit/tunnel flag, pin source fields, corpus repo, health URL) without the remove+add that drops the overlay, secret bindings, and deploy log. The verb the #1 VM rename wanted (session 0022 had to edit the record by hand). 5. **`flotilla deploy reconcile [--json]`.** Probes `/api/health` once and settles the log against reality: a stuck `in_progress` row is finalized `succeeded` when health confirms the target version (releasing the lock, without re-running the deploy); an already-terminal row is never rewritten — reconcile records a fresh health snapshot and reports the truth (including the "recorded `failed` but health says target is up" case the pre-drain fix makes rare). Exits non-zero when not confirmed healthy. SPEC folded back (§19.3 rule 2): §8.1 step 7 rewritten for the pre-drain, §12.1 verb inventory gains the two verbs, and a new §19.2 candidate ("graceful app-cooperative SSE drain" — a framework-side drain endpoint in rfc-app that would make the SIGKILL fallback go unused) records the remaining future work. **Tests:** 180 green (157 baseline + 23 new) in the isolated worktree venv. ## Decisions - **Scope = deploy-robustness, not GCP-resource alignment.** flotilla v1.2.0 is referenced in roadmap #34 as a possible slot, but #34 is the GCP resource-rename arc. This is a different, disjoint scope; #34 is untouched and unclaimed. Did NOT edit `ohm-rfc/ROADMAP.md` (no rfc-app/OHM surface moved, and 0026/0027 may be editing it). - **Pre-drain is flotilla-only.** A graceful SIGTERM exit needs app cooperation (rfc-app closing SSE streams on a signal); that's cross-repo and stays a §19.2 candidate. The SIGKILL fallback is the flotilla-only mitigation that needs no framework change and makes the v0.6.0 fault un-reproducible from flotilla's side today. - **reconcile never rewrites terminal rows.** Finalizing a stuck `in_progress` is honest (the deploy did complete; only the terminal write was lost). Rewriting a `failed`/`aborted` row to `succeeded` would erase the audit trail, so those get a health snapshot + an advisory instead. - **Branches+PR, no direct main push, no tag.** Per the operator directive and the sole-merger rule, I pushed the branch and left the merge + tag to the operator. ## Friction (continued) - **No gitea PR-create automation available.** `tea` CLI not installed; no gitea API token in env (HTTPS creds live in macOS keychain — not fished out). So the PR is "pushed + create-URL handed to operator," not auto-opened. The operator opens + merges anyway. ## Operating notes / cleanup - Isolated worktree left in place at `/Users/benstull/projects/wiggleverse/ohm-rfc-app-flotilla-s0028` (its own `.venv`) in case the operator wants to iterate on the PR. Remove with `git worktree remove ohm-rfc-app-flotilla-s0028` after the PR merges. - The canonical dir `…/ohm-rfc-app-flotilla` was restored to `main` carrying the OTHER session's (0026/0027) unstaged secret-redaction edits, untouched — that work is theirs to commit on their branch. --- ## Handoff prompt for the next session (Delivered to the operator in chat before this transcript was published.) > You are an OHM roadmap driver session, running in parallel with others. > CLAIM YOUR SESSION ID FIRST: > `~/git/ohm-infra/scripts/claim-session-id.sh --start ` > It prints the active in-progress sessions — coordinate if your work overlaps. > Read `~/git/ohm-infra/SESSION-PROTOCOL.md` §7 + PULL `ohm-rfc/ROADMAP.md` once. > > STATE (end of 0028, 2026-05-28): OHM live at rfc-app v0.24.0 on the renamed > VM `ohm-rfc-app`; deploys unblocked, lock free. flotilla **v1.2.0** is built > + pushed as a **PR** (deploy robustness + `deployment update` + `deploy > reconcile` + phase-7 SSE pre-drain) — NOT merged. Operator: open + merge > + tag `v1.2.0`: > https://git.wiggleverse.org/wiggleverse/ohm-rfc-app-flotilla/pulls/new/session-0028/flotilla-v1.2.0 > Note: v1.2.0 and the concurrent secret-redaction work (sessions 0026/0027, > if they ship it) both touch `deploy.py` `_PhaseRunner` → expect a small > merge conflict on the second PR; resolvable (distinct concerns). > > **CRITICAL for parallel flotilla work:** branches alone do NOT isolate > concurrent sessions sharing one working dir — take a **manual git worktree** > (`git worktree add -b main`) with its own venv > (`python3 -m venv .venv && .venv/bin/pip install -e .[test]`; the canonical > `.venv`'s editable finder is pinned to the shared dir). > > PICK ONE disjoint item. Open candidates: #28 (PR cross-references — large, > its own session); #21 Part A Amplitude audit (~week of real data, > ~2026-06-04); #22 pro-consent copy (needs operator-drafted + counsel text); > #20/#18 ops (DMARC Phase B ~2026-06-04, stale `wiggleverse/meta` webhook > still pointing at deprovisioned `rfc.wiggleverse.org`, bounce wiring); > #31b/#25-redesign (need operator screenshots); #33/#34 repo + GCP name > alignment (operator-led). AVOID shipped surfaces (#26 ProposeModal/PR/RFC > views + `proposed_use_cases`; #27 suggest-tags; #29 sign-in resume + > `/api/auth/me` + `lib/useLastState.js`; v0.21.0 polish surfaces). > > Hard rules: never ask for secret bytes (give the `pbpaste | flotilla secret > set ohm-rfc-app ` gesture); coordinate the serialized deploy lock + > version slot with the operator; subagents write their own §5 subsession > transcripts; finalize + publish your transcript LAST (deliver the handoff > prompt before publish). Use branches + PRs. Claim your ID, then go.