Files
octopus-workflow/core/skills/headless-session-ops/reference/lifecycle-ops.md
T

107 lines
6.1 KiB
Markdown
Raw Normal View History

> Extracted from headless-session-ops/SKILL.md (Post-launch lifecycle: health check, hung rescue, stale-worldview correction ([org-internal #2459])) — moved verbatim 2026-08-26, ticket [org-internal #3480].
## Post-launch lifecycle: health check, hung rescue, stale-worldview correction ([org-internal #2459])
Launching is half the job — a headless worker can hang silently or wake up
with an outdated worldview. All three failure modes below were observed in
production on 2026-08-16 during Epic [org-internal #2459] (ops-restart window); the runbooks
are extracted from that incident record.
### Health check — three states, one poll pattern
Poll two signals together — `GET /session?directory=` for `.state` (plus the
per-session `progress` object when present), and `GET
/session/:id/message?limit=1` for the newest message's `.info.id` — **twice,
spaced 3060 s**. `MessageID` is monotonic (`MessageID.ascending`,
`message.ts`), so a frozen newest-id across both samples is exactly the old
"message count frozen" signal at O(1) per poll instead of O(transcript) —
never poll the no-`limit` form for liveness, it loads the whole transcript
server-side:
| Verdict | Signals | Action |
|---|---|---|
| healthy | `state=generating` AND (`progress.stepCount` growing OR newest message id advancing) | leave it alone |
| idle | `state=idle` AND `progress` absent-or-stale across both polls | turn ended — read the LAST message (same `limit=1` fetch): task unfinished → wake prompt; `status=done` report → harvest |
| **hung** | `state=generating` AND `progress.stepCount`/`lastStepAt` AND newest message id ALL frozen across both polls | the generation stream is dead; queued prompts will NEVER be consumed — rescue required |
`state` alone lies in BOTH directions ([org-internal #3215]): a hung session still reports
`generating` (field sample: session A, 2026-08-16 — `state=generating` for 6+
minutes with the message count frozen at 478), and a healthy mid-step worker
can read `idle`/empty message tails in a single snapshot (2026-08-23 W3 wave:
4 healthy workers aborted off one snapshot). **The dual-sample delta is the
discriminator; a single snapshot is NEVER an abort basis.**
`progress` (`stepCount` cumulative LLM-round counter, `lastStepAt` epoch-ms
heartbeat — [org-internal #3215]) moves on every round even when message tails are
transiently empty or `state` flickers; it is absent for sessions that never
ran since instance start (treat absent = no signal, fall back to count
deltas). `lastStepAt` freshness alone does NOT prove liveness (a long tool
call inside one round keeps it stale for minutes) — always compare TWO
samples spaced ≥30 s.
### Hung rescue — abort, then re-wake (in this order)
```sh
# 1. Abort the dead stream (queued-but-unconsumed prompts do NOT unblock it)
curl -s -X POST "$BASE/session/$SID/abort?directory=$DIR" -o /dev/null -w "%{http_code}\n" # → 200
# 2. Verify idle
curl -s "$BASE/session?directory=$DIR" | jq -r --arg s "$SID" '.[] | select(.id==$s) | .state' # → idle
# 3. Check the worktree — the hung turn may have left uncommitted files (NOT lost)
git -C <worktree> status --short
# 4. Re-deliver the wake prompt (same Iron Law: known-good model).
# A prompt queued BEFORE the abort may still never fire — always re-send.
```
Endpoint: `POST /session/:sessionID/abort` (`groups/session.ts:113`,
`handlers/session.ts:311`). Abort stops the run loop; filesystem writes the
hung turn already made survive — inspect the worktree and list any recovered
files IN the wake prompt so the worker re-validates them instead of redoing
work (field sample: session A's hung turn had produced 3 src + 1 test file
that its revived self adopted).
**Anti-re-hang clause — include in every wake prompt.** The trigger for the
observed hang was a system-injected "请在适当的时机压缩当前会话" (compact at an
appropriate time): the worker ended its turn after compacting, leaving the
task half-done and idle. A wake prompt MUST carry, verbatim:
```text
若系统再注入「请在适当的时机压缩当前会话」:执行压缩后立即在后续 turn 继续任务,
绝不在任务未完成时以 idle 结束。
```
### Stale worldview — fact-baseline injection (correct BEFORE it acts)
A worker woken after an ops restart / long idle carries the worldview it went
to sleep with. It may re-dispatch superseded work, overwrite newer state, or
claim authority it does not hold (field samples, 2026-08-16: a revived worker
re-ran an M-01 verify another session had already published as FAIL —
deduplicated via flag #comment-21590; an ops-notification session inherited an
"orchestrator" identity from a compaction summary and announced a takeover —
corrected via 勘误 #comment-21654/[org-internal #21668]).
When you detect a stale-worldview session, do not wait for it to finish being
wrong — inject a fact-baseline prompt immediately (regular `prompt_async`),
structured as:
1. **You were woken; your worldview is stale** — name the event (restart /
maintenance window) and the current time.
2. **Authoritative state** — numbered facts with artifact links (wiki page,
issue comment), each with its timestamp; state explicitly which of the
recipient's standing assumptions are now INVALID.
3. **Your actual assignment now** — one concrete task (or explicit standby).
4. **Evidence rule** — verify each fact at its cited source before acting;
never act on this baseline alone.
(Field sample: the 2026-08-16 injection to session B pivoted it from the
superseded M-01 verify to the N-04b fix within one turn — the format works.)
### Authorization asymmetry — read this BEFORE "correcting" anyone
Only the session that CURRENTLY holds the authority may inject a baseline or
re-task a worker. If YOU might be the stale one — you woke from a restart,
your context came from a compaction summary, you cannot find your claim in
the durable record — assume YOU are stale: verify your identity/authority
against the record (issue assignee, claim comment, orchestrator session id)
BEFORE issuing any instruction. See the identity-verification clause in
`core/rules/compact.md` (recovery contract).