From 92738310e21a9937324eaff3abb25688d0c6e4fb Mon Sep 17 00:00:00 2001 From: Ben Stull Date: Thu, 28 May 2026 14:32:45 -0700 Subject: [PATCH] add 0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--2026-05-28T14-32.md + replace placeholder/variant SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--INPROGRESS.md --- ...RIPT-2026-05-28T14-05--2026-05-28T14-32.md | 228 ++++++++++++++++++ ...TRANSCRIPT-2026-05-28T14-05--INPROGRESS.md | 15 -- 2 files changed, 228 insertions(+), 15 deletions(-) create mode 100644 0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--2026-05-28T14-32.md delete mode 100644 0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--INPROGRESS.md diff --git a/0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--2026-05-28T14-32.md b/0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--2026-05-28T14-32.md new file mode 100644 index 0000000..7113465 --- /dev/null +++ b/0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--2026-05-28T14-32.md @@ -0,0 +1,228 @@ +# Session 0028.0 — Transcript + +> Date: 2026-05-28 +> Start: 2026-05-28T14-05 (PST implied) +> End: 2026-05-28T14-32 (PST implied) +> Subject: flotilla v1.2.0 — deploy robustness + record-maintenance verbs (PR, not merged) +> Status: **FINALIZED.** + +--- + +## Plan + +**Item picked:** flotilla **v1.2.0** — deploy-robustness + record-maintenance +verbs (the §19.2 infra candidate surfaced by Session 0022). Pure-flotilla, +touches no rfc-app code → zero collision with the in-flight rfc-app feature +sessions 0026/0027. No OHM deploy, no version-slot contention, no secret bytes. + +Working on branch `session-0028/flotilla-v1.2.0`; shipping as a **PR** (operator +merges + tags — per the "let's use branches and PRs" directive this session). + +Deliverables: +1. **Preflight SSH-failure disambiguation** (deploy.py phase 2). Today any + non-zero `test -d ` is reported "install dir missing" — but a + `gcloud compute ssh` connection/auth/IAP failure also returns non-zero, + so a connectivity failure is misreported as a missing dir (deploy.py:355). + Fix: a connectivity probe runs first; its failure reports the real cause. +2. **SshError → clean phase failure.** `_PhaseRunner.run` only catches + `_PhaseFailure`; an `ssh.SshError` (e.g. a restart SSH timeout) propagates + and leaves the `deploys` row stuck `in_progress` holding the §8.3 lock. + Catch it, record a phase failure, finalize the row. +3. **SSE pre-drain restart** (phase 7). Replace `systemctl restart` with a + flotilla-controlled stop (SIGTERM → bounded drain window → SIGKILL fallback + for lingering SSE keepalive streams) → start. Bounds the window flotilla + waits regardless of the unit's TimeoutStopSec. This is the v0.6.0 (Session + 0009/"I") failure made un-reproducible from the flotilla side. A graceful + app-cooperative drain endpoint in rfc-app stays a future §19.2 candidate. +4. **`flotilla deployment update`** verb (registry + CLI) — update mutable + record fields without remove+add. Motivated by the #1 VM rename. +5. **`flotilla deploy reconcile`** verb — probe `/api/health` and reconcile the + deploy log to reality: finalize a stuck `in_progress` row as `succeeded` + when health confirms the target version (releases the lock); honestly report + terminal rows health contradicts without falsifying the audit trail. + +Each lands with unit tests (no network — the deploy seams are DI'd). CHANGELOG +1.2.0 + VERSION/pyproject bump. Transcript finalized + published LAST. + +Parallel context at claim: sessions 0026 + 0027 also in-progress (launch +prompts not captured in their placeholders). flotilla-only scope chosen +specifically to stay disjoint from whatever rfc-app surfaces they touch. + +--- + +## Friction log + +- **14:09 — shared-working-directory collision.** Branches alone don't + isolate concurrent sessions: git tracks one working tree per checkout. + Shortly after I branched in the canonical flotilla dir, another live + session (0026 or 0027) wrote `ohm_rfc_app_flotilla/deploy.py` + + `tests/test_deploy.py` in the **same** directory (a secret-redaction + hardening, "finding L1"). Their files overlap exactly with mine + (`_PhaseRunner`, phase details). Editing in place would corrupt both + sessions' work. Resolution: restored the shared dir to `main` with their + unstaged edits preserved, then created a **manual git worktree** at + `…/ohm-rfc-app-flotilla-s0028` on my branch with its own fresh venv + (`pip install -e .[test]`; the canonical `.venv`'s PEP 660 editable + finder is pinned to the shared dir, so a shared venv would defeat + isolation). Worktree baseline: 157 tests green on clean `c26f251`. + **Handoff implication:** parallel flotilla sessions MUST each take a + worktree; the operator's "branches and PRs" guidance needs the worktree + step to actually isolate. Their redaction change and my v1.2.0 change + both touch `deploy.py` → expect a merge conflict in `_PhaseRunner` when + the second PR merges; resolvable (different concerns: they redact phase + detail, I add an `SshError` branch + DI redactor coexist cleanly). + +## Launch prompt + +You are an OHM roadmap driver session, running in parallel with others. +Claim session ID first. STATE (end of 0022): OHM live at rfc-app v0.24.0 on +the renamed VM `ohm-rfc-app`; deploys unblocked. Pick ONE disjoint roadmap +item. Strong infra candidate: flotilla v1.2.0 — fix the preflight that +misreports an SSH/auth failure as "install dir missing" (deploy.py:355), add +a `deployment update` verb, add `flotilla deploy reconcile` + SSE pre-drain +(§19.2 from this session). "Let's use branches and PR's so we don't need to +worry about conflicting with other sessions." + +--- + +## What shipped + +flotilla **v1.2.0**, on branch `session-0028/flotilla-v1.2.0`, pushed to +`origin` (gitea). **Open as a PR (not yet merged):** + + https://git.wiggleverse.org/wiggleverse/ohm-rfc-app-flotilla/pulls/new/session-0028/flotilla-v1.2.0 + +Commit `51c7ea8`. No OHM deploy, no rfc-app release, no pin bump — this is +operator-tool code. Operator opens + merges the PR (sole-merger rule) and +tags `v1.2.0` after merge. + +Five changes, all flotilla-internal, each unit-tested (the deploy seams are +dependency-injected so nothing touches gcloud or the network): + +1. **Preflight SSH-failure disambiguation** (`deploy.py` phase 2). The + misreport at `deploy.py:355`: a bare `test -d ` whose any-non-zero + return was read as "install dir missing" — but `gcloud compute ssh` exits + non-zero on a connection/auth/IAP failure too, so an unreachable VM was + misdiagnosed as a missing dir. Fix: a trivial `true` connectivity probe + runs first; its non-zero return is reported as the auth/connectivity + failure it is. (Initially gold-plated with a stdout marker; dropped that + for returncode-only after it would have forced every test stub — and every + future one — to simulate the echo. Returncode-only fully solves the actual + bug.) + +2. **`_PhaseRunner` catches `ssh.SshError`** → records a clean `failed` phase + instead of unwinding past `run_deploy` and stranding the row `in_progress` + holding the §8.3 lock. This is the structural cause of the v0.6.0 / Session + 0009 "aborted but actually fine" — a restart SSH timeout raised mid-phase. + +3. **Phase-7 SSE pre-drain.** Replaced `systemctl restart` with a + flotilla-controlled `systemctl kill -s TERM` → bounded `drain_wait_seconds` + (default 20s) → `systemctl kill -s KILL` for lingering SSE keepalive + streams → `systemctl start`. Bounded by `drain + restart_wait` regardless + of the unit's `TimeoutStopSec`, so a held-open SSE stream can't blow past + the SSH timeout. The phase detail records whether it drained via graceful + SIGTERM or had to SIGKILL. + +4. **`flotilla deployment update [--flag …]`** (registry + + CLI). In-place edit of mutable record fields (VM name/zone/project/service + user/install dir/systemd unit/tunnel flag, pin source fields, corpus repo, + health URL) without the remove+add that drops the overlay, secret + bindings, and deploy log. The verb the #1 VM rename wanted (session 0022 + had to edit the record by hand). + +5. **`flotilla deploy reconcile [--json]`.** Probes `/api/health` once + and settles the log against reality: a stuck `in_progress` row is finalized + `succeeded` when health confirms the target version (releasing the lock, + without re-running the deploy); an already-terminal row is never rewritten + — reconcile records a fresh health snapshot and reports the truth + (including the "recorded `failed` but health says target is up" case the + pre-drain fix makes rare). Exits non-zero when not confirmed healthy. + +SPEC folded back (§19.3 rule 2): §8.1 step 7 rewritten for the pre-drain, +§12.1 verb inventory gains the two verbs, and a new §19.2 candidate +("graceful app-cooperative SSE drain" — a framework-side drain endpoint in +rfc-app that would make the SIGKILL fallback go unused) records the remaining +future work. + +**Tests:** 180 green (157 baseline + 23 new) in the isolated worktree venv. + +## Decisions + +- **Scope = deploy-robustness, not GCP-resource alignment.** flotilla v1.2.0 + is referenced in roadmap #34 as a possible slot, but #34 is the GCP + resource-rename arc. This is a different, disjoint scope; #34 is untouched + and unclaimed. Did NOT edit `ohm-rfc/ROADMAP.md` (no rfc-app/OHM surface + moved, and 0026/0027 may be editing it). +- **Pre-drain is flotilla-only.** A graceful SIGTERM exit needs app + cooperation (rfc-app closing SSE streams on a signal); that's cross-repo + and stays a §19.2 candidate. The SIGKILL fallback is the flotilla-only + mitigation that needs no framework change and makes the v0.6.0 fault + un-reproducible from flotilla's side today. +- **reconcile never rewrites terminal rows.** Finalizing a stuck + `in_progress` is honest (the deploy did complete; only the terminal write + was lost). Rewriting a `failed`/`aborted` row to `succeeded` would erase + the audit trail, so those get a health snapshot + an advisory instead. +- **Branches+PR, no direct main push, no tag.** Per the operator directive + and the sole-merger rule, I pushed the branch and left the merge + tag to + the operator. + +## Friction (continued) + +- **No gitea PR-create automation available.** `tea` CLI not installed; no + gitea API token in env (HTTPS creds live in macOS keychain — not fished + out). So the PR is "pushed + create-URL handed to operator," not + auto-opened. The operator opens + merges anyway. + +## Operating notes / cleanup + +- Isolated worktree left in place at + `/Users/benstull/projects/wiggleverse/ohm-rfc-app-flotilla-s0028` (its own + `.venv`) in case the operator wants to iterate on the PR. Remove with + `git worktree remove ohm-rfc-app-flotilla-s0028` after the PR merges. +- The canonical dir `…/ohm-rfc-app-flotilla` was restored to `main` carrying + the OTHER session's (0026/0027) unstaged secret-redaction edits, untouched + — that work is theirs to commit on their branch. + +--- + +## Handoff prompt for the next session + +(Delivered to the operator in chat before this transcript was published.) + +> You are an OHM roadmap driver session, running in parallel with others. +> CLAIM YOUR SESSION ID FIRST: +> `~/git/ohm-infra/scripts/claim-session-id.sh --start ` +> It prints the active in-progress sessions — coordinate if your work overlaps. +> Read `~/git/ohm-infra/SESSION-PROTOCOL.md` §7 + PULL `ohm-rfc/ROADMAP.md` once. +> +> STATE (end of 0028, 2026-05-28): OHM live at rfc-app v0.24.0 on the renamed +> VM `ohm-rfc-app`; deploys unblocked, lock free. flotilla **v1.2.0** is built +> + pushed as a **PR** (deploy robustness + `deployment update` + `deploy +> reconcile` + phase-7 SSE pre-drain) — NOT merged. Operator: open + merge +> + tag `v1.2.0`: +> https://git.wiggleverse.org/wiggleverse/ohm-rfc-app-flotilla/pulls/new/session-0028/flotilla-v1.2.0 +> Note: v1.2.0 and the concurrent secret-redaction work (sessions 0026/0027, +> if they ship it) both touch `deploy.py` `_PhaseRunner` → expect a small +> merge conflict on the second PR; resolvable (distinct concerns). +> +> **CRITICAL for parallel flotilla work:** branches alone do NOT isolate +> concurrent sessions sharing one working dir — take a **manual git worktree** +> (`git worktree add -b main`) with its own venv +> (`python3 -m venv .venv && .venv/bin/pip install -e .[test]`; the canonical +> `.venv`'s editable finder is pinned to the shared dir). +> +> PICK ONE disjoint item. Open candidates: #28 (PR cross-references — large, +> its own session); #21 Part A Amplitude audit (~week of real data, +> ~2026-06-04); #22 pro-consent copy (needs operator-drafted + counsel text); +> #20/#18 ops (DMARC Phase B ~2026-06-04, stale `wiggleverse/meta` webhook +> still pointing at deprovisioned `rfc.wiggleverse.org`, bounce wiring); +> #31b/#25-redesign (need operator screenshots); #33/#34 repo + GCP name +> alignment (operator-led). AVOID shipped surfaces (#26 ProposeModal/PR/RFC +> views + `proposed_use_cases`; #27 suggest-tags; #29 sign-in resume + +> `/api/auth/me` + `lib/useLastState.js`; v0.21.0 polish surfaces). +> +> Hard rules: never ask for secret bytes (give the `pbpaste | flotilla secret +> set ohm-rfc-app ` gesture); coordinate the serialized deploy lock + +> version slot with the operator; subagents write their own §5 subsession +> transcripts; finalize + publish your transcript LAST (deliver the handoff +> prompt before publish). Use branches + PRs. Claim your ID, then go. diff --git a/0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--INPROGRESS.md b/0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--INPROGRESS.md deleted file mode 100644 index adc9df7..0000000 --- a/0028/SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--INPROGRESS.md +++ /dev/null @@ -1,15 +0,0 @@ -# Session 0028.0 — Transcript - -> Date: 2026-05-28 -> Start: 2026-05-28T14-05 (PST implied) -> Status: **PLACEHOLDER — claimed at session start; finalized at session end.** -> -> This file reserves session ID 0028. The driver replaces this body -> with the full transcript before publishing, and renames the file to -> its final SESSION-0028.0-TRANSCRIPT-2026-05-28T14-05--.md form. - ---- - -## Launch prompt - -_(launch prompt not captured at claim time)_