Initial publish v0.1.0: standalone workflow core (corpus + examples + guards)

This commit is contained in:
octopus
2026-09-15 08:41:51 +08:00
commit bb35e661b2
114 changed files with 20240 additions and 0 deletions
+547
View File
@@ -0,0 +1,547 @@
---
name: headless-session-ops
description: >
Use ONLY when an agent must autonomously launch a headless main session 鈥?
create a session and drive agent generation over the HTTP `prompt_async`
endpoint with NO persistent client attached (Web UI / TUI absent) 鈥?OR
lifecycle-manage such a session afterwards: health-check polling, hung-stream
diagnosis (state=generating + frozen message count), abort + re-wake rescue,
provider quota-wall recovery (worker killed mid-task by a usage limit 鈥?
detect, re-wake on a known-good provider, inject facts), and fact-baseline
injection to correct a stale worldview in a woken worker.
This is the path used by in-session agents spawning sibling sessions,
`octopus run --attach` non-interactive mode, and cron / CI drivers. Covers
the 3-step flow (create session 鈫?POST prompt_async WITH a known-good
`model` 鈫?poll `/session/:id/message` to verify generation), how to REUSE
the current session's model for the new session, the model-field
requirement, the directory default (current session's directory unless
targeting another repo), and the ticket-driven recipe (launch a session
seeded from issue #N). Do NOT use when a persistent client (Web UI / TUI)
is driving the session 鈥?those use the synchronous `prompt` endpoint and
always send `model`.
triggers:
# English 鈥?the brand phrases for headless launch.
- headless session
- launch session
- prompt_async
- prompt async
- autonomous session
# English 鈥?the failure mode (agent sees a stalled session and needs to know why).
- session not generating
- session generation stuck
# English 鈥?ticket-driven launch.
- launch session for issue
- launch session for ticket
- spawn session for ticket
# English 鈥?patrol mode (delayed self-wake polling loops, [org-internal #3937]).
- patrol loop
- patrol mode
- delayed wake
# Chinese 鈥?bare noun phrases (matches() is a contiguous substring).
- 鎷夎捣涓讳細璇?
- 鎷夎捣浼氳瘽
- 鑷不浼氳瘽
- 鏃犲ご浼氳瘽
# Chinese 鈥?patrol mode ([org-internal #3937]).
- 宸℃妯″紡
- 寤惰繜鍞ら啋
- 浼氳瘽涓嶇敓鎴?
- 浼氳瘽鍗℃
- 涓哄伐鍗曟媺璧蜂細璇?
- 宸ュ崟鎷夎捣浼氳瘽
# English 鈥?lifecycle (hung rescue, stale-worldview correction, [org-internal #2459]).
- abort session
- revive session
- session rescue
- fact baseline
# English 鈥?quota-wall recovery ([org-internal #3669]): worker killed by provider usage limit.
- quota wall
- usage limit reached
- quota exhausted
- switch provider
# Chinese 鈥?lifecycle ops ([org-internal #2459]).
- 鍞ら啋浼氳瘽
- 浼氳瘽鎸傛
- 浜嬪疄鍩虹嚎
# Chinese 鈥?quota-wall recovery ([org-internal #3669]).
- 閰嶉澧?
- 閰嶉鑰楀敖
- 鐢ㄩ噺涓婇檺
- 鎹緵搴斿晢
role: Producer
---
> Core 涓珛鐗堬紙Increment 6a 鏀瑰啓锛屽師 deferHard verbatimDir锛夈€傛満鍒躲€佺粨鏋勪笌 frontmatter 淇濇寔锛涘疄渚嬫湳璇紙宸ュ叿鍚嶃€佽矾寰勩€佸伐鍗曞彿锛夋寜 `core/adapters/TERMINOLOGY.md` 缁戝畾鍒板叿浣撳疄渚嬨€?
# Headless Main Session Launch (prompt_async)
Launch a **main session for a specific issue / work item from a headless
context** 鈥?no Web UI, no TUI, just HTTP. This is the "fire-and-forget" path: an
agent inside another session, a cron job, a CI script, or `octopus run --attach`
non-interactive mode all land here when they need to spin up a session that
actually generates.
## Directory convention 鈥?independent session in the current workspace, no worktree
The default is an **independent session sharing the current session's
directory** 鈥?the repo you are already working in. No new worktree, no
checkout, no extra `bun install` or `.codegraph/` rebuild. Session creation
and worktree discipline are **orthogonal**: this skill only opens a session
record pointed at a directory; whether the launched session later creates its
own worktree is a _workflow_ decision (per the worktree-discipline rule) it
makes when it starts actual work 鈥?not something the launcher does. One
exception: the **ticket-driven recipe** below, where claim-first ([org-internal #2297])
requires the launcher to push the workflow branch before launch, making a
pre-built worktree the claim carrier + ready workspace (see Pitfall #10
carve-out). So, unless a different repo / path is specifically targeted,
`?directory=` defaults to the current session's directory and is NOT mandatory:
- **You are an in-session agent** (most common): set `DIR` to your own
session's `directory` (you already know it from context, or read `.directory`
from your own session record via `GET /session/$LAUNCHER_SID`). The new
session runs independently in that same workspace. Pass an explicit
different `?directory=` only when the new session must work in another
**repo** 鈥?never a worktree of the same repo. The Web UI sidebar groups
sessions by directory (`sidebar-project.tsx` per-workspace
`workspaceSessions(directory)`, fetched via `session.list({ directory })`),
so a session pointed at a worktree disappears from the workspace the creator
manages. A pre-built ticket worktree is handed to the launched session via
the claim comment + seed prompt, NOT via `?directory=`.
- **No current session** (cron / CI / bare script): the server-side fallback
when `?directory=` is omitted is the **server process's `process.cwd()`**
(`server/routes/instance/httpapi/middleware/workspace-routing.ts:76`), which
is wherever the server was started 鈥?unreliable. In that context `?directory=`
remains effectively required.
- **Dispatch convention 鈥?directory = code location, not tracker repo.** When
launching a session for a ticket, set `?directory=` to the **main workspace**
checkout that holds the **code the ticket edits**, which may differ from the
repo where the ticket is tracked. Example: `<owner>/<backend-repo>#<n>` is filed in
the ticket-tracker repo but edits the main repo's workflow files 鈫?the
session's `?directory=` is the **code checkout root**
(`<workspace-root>/<org>/<repo>`), not the tracker repo checkout. Never point
`?directory=` at a ticket worktree 鈥?the worktree goes in the claim comment +
seed prompt, not the directory field (sidebar visibility, see bullet 1). The
owner-check (`GET /session?directory=$DIR`) is scoped per directory, so
rooting the session at the correct code checkout is also what makes the
duplicate-session guard meaningful. Server-side backstop ([org-internal #3190]): creating a
session whose directory IS a linked git worktree answers with an
`X-Session-Directory-Warning` response header and a server warn log naming
the owning main repo 鈥?if you see that header, the session will be invisible
to `GET /session?directory=<mainRepo>` owner-checks; re-check your launch
parameters before proceeding.
> **Authoritative reference:** `rules/headless-session-ops` wiki page
> (<<instance-base-url>/Octopus/octopus/wiki/rules%2Fheadless-session-ops>).
> This skill is its runtime carrier. Provenance: issue **[org-internal #1695]** (which
> corrected the misdiagnosed [org-internal #1691], see comment 9205). Code citations below were
> verified against `<harness-package>` at HEAD.
## The `model` field 鈥?always provide a known-good model (read this first)
`prompt_async` is **fire-and-forget**: after HTTP returns 204 (or 202 when the
message is queued behind a wedge/zombie session 鈥?no live runner in that
process), the server runs the full agent loop in a separate fiber. **Always
include an explicit `model`
in the payload, and reuse the current session's model for it** (see the next
section). Two reasons:
1. **Current code does not hard-require `model`.** The HTTP boundary marks
`model` as optional (`<harness-package>/src/session/prompt.ts:1860`), and
`createUserMessage` fills a missing model via a fallback chain
(`prompt.ts:858`):
```ts
const model = input.model ?? ag.model ?? yield * currentModel(input.sessionID)
// ^^^^^^^^^ agent ^^^^^^^^^^^^^^^^^ session/default
```
So omission does **not** produce a clean validation error 鈥?it silently
resolves to the agent's model, the session's last-used model, or the
provider default (`currentModel`, `prompt.ts:824-838`).
2. **But the fallback is not safe in the autonomous context.** If it resolves
to a model that is unusable here (region-blocked, wrong subscription tier,
not configured), the runLoop's `getModel` raises `ModelNotFoundError`
(`prompt.ts:1399`, `:804-822`). In the **async** handler that error is caught
by the delivery fork's failure handling (`forkPromptDelivery` in
`handlers/session.ts`) and never reaches the HTTP
caller 鈥?so the session looks healthy (HTTP 204, user message persisted) yet
**never generates**. This is exactly the "silent stall" observed empirically
in [org-internal #1695]: a session with a 1-message transcript (user only) and frozen
`time.updated`.
| payload | result |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| `{agent, parts}` 鈥?no `model` | 204/202; model filled by fallback; if the resolved model is unusable 鈫?**apparent silent stall** (no error surfaced, no generation). |
| `{agent, model:{providerID,modelID}, parts}` | 204/202; deterministic generation within seconds (verified: "1+1 绛変簬 2銆? in ~6s). |
> **Iron Law:** every `prompt_async` payload MUST carry a **known-good** `model`,
> and the safest known-good model is **the one driving this session right now**
> (it is, by definition, generating). Never rely on the fallback chain in a
> headless/autonomous launch.
**Field-name gotcha when copying a model:** the prompt payload's `ModelRef` uses
`{ providerID, modelID }` (`prompt.ts:1852-1855`), but the **session record's**
model uses `{ id, providerID, variant }` (`session.ts:256-260`) 鈥?i.e. `modelID`
is called `id`. When you reuse a model read from a session record, remap
`model.id 鈫?modelID` (see recipe below). User-message `info.model` already uses
`modelID` (`message.ts:403-407`).
## Reuse the current session's model (recommended default)
The model powering the launching session is proven to work 鈥?reuse it. Discover
it by whichever path fits your context:
- **You are an in-session agent** (most common): read it straight from your own
system prompt, which states verbatim:
> "You are powered by the model named **{modelID}** (provider: **{providerID}**)."
e.g. `modelID = "glm-5.2"`, `providerID = "zai-coding-plan"`锛堟浘鍚?
zhipuai-coding-plan锛?026-08 鏇村悕锛?
- **Programmatic / no system prompt** (cron, CI, sibling caller): `GET` the
launching session's record and read `.model`:
```sh
curl -s "$BASE/session/$LAUNCHER_SID?directory=$DIR" \
| jq '.model' # 鈫?{ "id": "glm-5.2", "providerID": "zai-coding-plan", "variant": ... }
```
Then **remap** `model.id 鈫?modelID` when building the payload.
- **Verify the provider is actually connected** before launching (cheap
insurance 鈥?the model only generates if its provider is live):
```sh
curl -s "$BASE/provider?directory=$DIR" | jq '.connected | index("<current providerID>")'
# 鈫?null = provider NOT connected (do NOT launch); 0..n = connected
```
or call the `list-models` tool. Note: model-level availability is only
knowable at runtime 鈥?a _configured_ model can still fail (subscription-tier
or region errors); `GET /config/providers` lists _configured_ models, not
usable ones. The safest pre-flight is "provider connected" + reusing the
current session's model, which is generating by definition. There is **no
`GET /model` HTTP endpoint** (it 404s); do not use it.
## Prerequisites
- A reachable octopus HTTP instance (e.g. `http://127.0.0.1:4096`).
- The target directory (passed as `?directory=`). Defaults to the current
session's directory for in-session launches 鈥?see the Directory convention
above; only specify a different one explicitly when targeting another repo.
- A known-good `model` (reuse the current session's 鈥?see above). Model IDs
must match the provider config in `<config-home>/octopus/octopus.jsonc`;
credentials live in the octopus account keyring (not env), so an isolated
HOME cannot reuse them.
## The 3-step flow
Set the base URL and directory once:
```sh
BASE=http://127.0.0.1:4096
DIR=/data/octopus
```
### Step 1 鈥?Create the session
```sh
SID=$(curl -s -X POST "$BASE/session?directory=$DIR" \
-H 'content-type: application/json' \
-d '{"title":"#N 鈥?<short description>","agent":"builder"}' | jq -r .id)
echo "session=$SID"
```
- `POST /session` returns `Session.Info` with `.id` (`groups/session.ts:207-211`).
- `?directory=` points at the repo root. For in-session launches it defaults to
the **current session's directory** (Directory convention above) 鈥?keep it
that way so the new session stays visible in the creator's sidebar; pass an
explicit path only when targeting a different repo (a worktree of the same
repo is NOT a directory target 鈥?hand it over via the seed prompt), and
always pass it from cron / CI where no current session exists (server falls
back to its own `process.cwd()` otherwise).
- `agent` is `builder` (the main agent). Sub-agents (explorer/worker) are
spawned within a session 鈥?never named here.
### Step 2 鈥?Deliver the seed prompt (CRITICAL: include a known-good `model`)
```sh
curl -s -X POST "$BASE/session/$SID/prompt_async?directory=$DIR" \
-H 'content-type: application/json' \
-d '{
"agent": "builder",
"model": { "providerID": "<current providerID>", "modelID": "<current modelID>" },
"parts": [ { "type": "text", "text": "<seed prompt: restate goal, constraints, acceptance criteria, cite wiki artifact paths>" } ]
}'
```
- Endpoint is `POST /session/:sessionID/prompt_async`; it answers **204, no
body** when processing starts now (idle session, live runner, noReply
injection, or waiting-question preemption), **202** when the message is
queued behind a session that is generating with no live runner in this
process (wedge/zombie shape 鈥?stored and re-driven on restart drain or the
next runner), **202** when `delay_sec` defers delivery to an in-memory
delayed wake ([org-internal #3937] A 鈥?see "宸℃妯″紡 / Patrol mode" below), and **409**
when the session is paused (message NOT stored) or the maintenance-mode
hold queue is full.
- **Global route alternative ([org-internal #4307]):** when the client only knows the
sessionID (no `?directory=` routing), `POST /prompt_async` (no path prefix)
resolves the instance from the app database by the body's `sessionID` and
delegates to the same delivery core 鈥?same 204/202/409 semantics, same
`delay_sec` contract, body is the session-scoped payload plus a required
`sessionID` field. Unknown sessionID 鈫?404 naming the session-scoped
route. This is the natural fit for the completion-report POST (fixed port,
no directory context).
- `delay_sec` (optional, integer 0..86400, else 400) schedules the prompt for
a one-shot delayed self-wake: the server holds the prompt in memory and
fires it exactly once after `delay_sec` seconds, only for an idle session
at intake. **Precedence ([org-internal #3937] review):** `delay_sec` applies only when the
intake would otherwise start processing now (immediate class). For the
queued-behind-wedge 202 member and under maintenance queue mode it is
stripped 鈥?retention semantics win (the server logs the drop; the delay is
NOT honored). Pending wakes are **in-memory only 鈥?a server restart abandons
them**; a patrol loop must therefore tolerate a missed wake and re-arm.
Never use `sleep <big>; curl ...` in a bash tool call to wait 鈥?that blocks
the turn with zero output (the [org-internal #3937] incident shape); the bash tool now
warns on bare sleeps > 120s (long-sleep guard, [org-internal #3937] C).
- `model` is `{ providerID, modelID }` or the shorthand string
`"providerID/modelID"` (split on the FIRST `/`; accepted since [org-internal #4307] 鈥?
e.g. `"zai-coding-plan/glm-5.2"`). **Do not omit it.** Fill the values from
your own session's model (see "Reuse the current session's model" above) 鈥?
never hardcode a specific model. A string without a `/` is rejected with
400 naming both accepted forms.
- When the seed prompt contains newlines/quotes, build the payload from a file
to avoid shell-escaping errors:
```sh
jq -n --rawfile p seed.txt '{agent:"builder",model:{providerID:"<current providerID>",modelID:"<current modelID>"},parts:[{type:"text",text:$p}]}'
```
### Step 3 鈥?Verify the session actually generated (do NOT assume success)
HTTP 204/202 鈮?generation success. Poll until an `assistant` message appears. The
endpoint is **`/session/:id/message` 鈥?SINGULAR**, and it returns a bare array
whose elements are `{ info, parts }` with `role` at `.info.role`
(`groups/session.ts:88,183-194`). Poll the **tail with `limit=1`** 鈥?the
no-`limit` form loads the ENTIRE transcript server-side (O(messages脳parts) per
poll; pitfall #14's 478-message session paid it on every health check), while
`limit=1` returns just the newest message (REQ-F-010 windowing, `message.ts`
`page()`):
```sh
# wait 5鈥?0s, then poll the TAIL 鈥?never the full transcript
curl -s "$BASE/session/$SID/message?directory=$DIR&limit=1" \
| jq '{latest_role: .[0].info.role, latest_id: .[0].info.id}'
```
Verdict: `latest_role == "assistant"` 鈫?the first generation completed (the
seed is `user`; the newest message flips to `assistant` only once the model
answered 鈥?an errored turn also persists an assistant message carrying the 鈿狅笍
usage-limit part, [org-internal #2912], so check the tail text before declaring success).
Stays `user` across 2鈥? polls spaced 10鈥?0 s 鈫?**apparent silent stall**;
almost certainly the model did not resolve to a usable one (fallback landed
badly, or you omitted `model`). Redeliver Step 2 with an explicit known-good
`model`. Need everything newer than a known point? Page forward with
`after=<cursor>&limit=N` and follow the `X-Next-Cursor` response header 鈥?
just never omit `limit`.
> 鈿狅笍 Do NOT write `/messages` (plural) 鈥?that path does not exist for listing.
> (`POST /session/:id/message`, same singular path but POST, is the synchronous
> prompt endpoint; don't confuse the two.)
## Launch a session for a specific ticket (#N)
The common case: an agent (or cron/CI) spins up a fresh main session to work a
tracked issue 鈥?seed prompt composed **from the issue body**, launched on the
current session's model. **Step-0 owner check is MANDATORY ([org-internal #1803])**: four data
sources (session-title scan, assignee/claim, open PRs, remote branches); any
live claim 鈫?ABORT the launch (`session-scope-guard.md` points here for that
pre-step). Full recipe 鈥?owner-check scan script with `X-Total-Count` /
`X-Has-More` pagination handling ([org-internal #3190]), claim-first atomic 3-step ([org-internal #2297],
+ one-command provisioning via `script/claim-provision.sh`, [org-internal #3642]),
seed-prompt MUSTs, backend duplicate-ticket hard guard ([org-internal #1989], on by default
[org-internal #2350]), launch + poll commands: `reference/ticket-recipe.md` (read BEFORE
creating any ticket-driven session).
## Completion-report protocol (瀹屾垚鍥炴姤鍗忚, [org-internal #2374])
`prompt_async` is fire-and-forget **for the launcher too**: without this
protocol the orchestrator has NO push channel and must poll forever ([org-internal #2366]).
The worker actively reports `status=done|blocked|handoff` back into the
orchestrator's session via `prompt_async`, with an issue-comment fallback.
Default report format is `branch=<ref>`-based 鈥?workers never open PRs
(TD-678/[org-internal #4425]); `pr=#N` appears only on `uncoordinated` self-opens.
Hard constraints: a headless worker MUST NOT call the `question` tool
([org-internal #2378] 鈥?directional/irreversible decisions are `status=blocked` reports,
then STOP), and MUST NOT self-merge its PR / self-close the issue / touch
`main` ([org-internal #2386] 鈥?keep the PR open; the orchestrator merges and closes). Full
protocol 鈥?orchestrator-side seed block (verbatim template), worker-side
steps, decision-authority layering: `reference/completion-report.md` (read
BEFORE composing a seed prompt for any worker you need to hear back from).
## Post-launch lifecycle: health check, hung rescue, stale-worldview correction ([org-internal #2459])
Launching is half the job 鈥?a headless worker can hang silently or wake with
a stale worldview. Health check = two polls 30鈥?0 s apart (`.state` + `progress`
+ message count); `state=generating` with ALL counters frozen across BOTH
samples = hung 鈫?abort FIRST (`POST /session/:id/abort`), then re-wake with a
known-good model + the anti-re-hang clause. Stale worldview 鈫?fact-baseline
injection (only the CURRENT authority injects; if YOU might be the stale one,
verify identity against the durable record first). Full runbooks 鈥?verdict
table + [org-internal #3215] dual-sample rule, abort + re-wake sequence, fact-baseline
structure, authorization asymmetry: `reference/lifecycle-ops.md` (read when
managing a launched session). Quota wall 鈥?the worker's turn killed by a
terminal usage-limit error (looks like silent completion: no report, ticket
stalled) 鈥?has its own runbook: detection signals (finish=error 鈿狅笍 tail,
provider quota markers), provider-switch re-wake, fact-baseline template,
pre-dispatch provider check: `reference/quota-wall-recovery.md` ([org-internal #3669], evidence
[org-internal #3627]; read when a worker goes silent mid-task).
## 宸℃妯″紡 / Patrol mode (delayed self-wake, [org-internal #3937])
A patrol loop periodically nudges a headless session without holding a
connection or burning a turn on a blocking sleep. Since [org-internal #3937] A the server
supports this natively via `delay_sec` on `prompt_async`.
```sh
curl -s -X POST "$BASE/session/$SID/prompt_async?directory=$DIR" \
-H 'content-type: application/json' \
-d '{
"agent": "builder",
"model": { "providerID": "<current providerID>", "modelID": "<current modelID>" },
"parts": [ { "type": "text", "text": "<patrol instruction>" } ],
"delay_sec": 3600
}'
```
Rules of the road:
- **202 + `delay_sec` 鈮?queued-behind-wedge 202.** Both are 202, but the
delayed wake holds the prompt in memory and delivers it exactly once when
the delay elapses (session must have been idle at intake; `0` equals
omission 鈫?immediate 204 path). If the session was NOT idle at intake, the
intake keeps the queued semantics and `delay_sec` is ignored (logged
server-side) 鈥?the same precedence as maintenance queue mode, where the
intake is retained and `delay_sec` never delays the flush.
- **Paused at wake 鈫?dropped, not stored.** The wake's `prompt()` rejects and
the server publishes an error event; the message is never stored. A patrol
driver should check session state before assuming delivery.
- **Restart abandons pending wakes.** They are in-memory server fibers 鈥?no
persistence, no drain. After any server restart the patrol driver owns
re-arming the next wake; treat a missed wake as expected, not as an error.
- **Latest wake wins, per session ([org-internal #4069]).** Arming `delay_sec` on a session
that already has a pending wake SUPERSEDES it: the previous countdown is
cancelled, only the latest wake fires. Re-arming each turn is the intended
patrol pattern 鈥?a superseded wake never delivers its prompt and publishes
no error. (Pre-[org-internal #4069] stacks both FIFO 鈥?on old binaries, re-arm only after
the previous wake fired.)
- **Never emulate a delay with the bash tool.** `sleep 900; curl ...` blocks
the turn with zero output and gets manually aborted as a hang (the [org-internal #3937]
incident). The bash tool now prepends a `<shell_warning>` on bare sleeps
beyond 120s ([org-internal #3937] C) 鈥?treat that warning as a redirect to this section.
## Do NOT use this skill when
- A **persistent client** (Web UI / interactive TUI) is attached 鈥?those use the
synchronous `prompt` endpoint (it blocks until generation completes) and send
`model` automatically. Using `prompt_async` there gains nothing and loses the
in-line response. This includes the rescue runbook: a hung session driven by
a persistent client is the client user's to handle (refresh / re-prompt from
the UI) 鈥?abort + re-wake targets headless / managed sessions only.
- You want the response **inline** in the same HTTP call 鈥?`prompt_async` is
fire-and-forget; use the synchronous `prompt` endpoint instead.
## Common pitfalls
| # | Pitfall | Fix |
| --- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 1 | Omitting `model` / relying on the fallback 鈫?apparent silent stall (most common) | Always send a known-good `model:{providerID,modelID}` 鈥?reuse the current session's |
| 2 | Listing endpoint written `/messages` (plural) 鈫?404 | Use `/session/:id/message` (singular); returns a bare array, role at `.info.role` |
| 3 | Copying the session record's model verbatim into the payload | Remap field names: session `.model.id` 鈫?payload `modelID` (session uses `id`, payload uses `modelID`) |
| 4 | Stale `octopus` binary rejects newer config schema (`Unrecognized key`) | Use the source CLI `bun run ./src/index.ts ...`; the config itself is valid |
| 5 | Missing `?directory=` query param | In-session launch: default to the current session's directory. Cron/CI (no current session): pass the repo root explicitly 鈥?server falls back to its own `process.cwd()` |
| 6 | Using synchronous `prompt` for headless launch | `prompt` blocks until done 鈥?wrong for fire-and-forget; use `prompt_async` |
| 7 | Shell-escaping errors in a multi-line seed prompt | Build the payload with `jq -n --rawfile` from a file |
| 8 | Verifying models via `GET /model` 鈥?no such endpoint (404) | Use `GET /provider` 鈫?`.connected` (provider live check); `list-models` tool also works. `GET /config/providers` lists _configured_ models only |
| 9 | Launching a session for a ticket another live session already owns 鈫?conflicting diffs, duplicate work ([org-internal #1744]/[org-internal #1753]) | Run the Step-0 owner check (`GET /session` title scan + worktree/branch/PR scan); ABORT and report on conflict |
| 10 | Creating a worktree / fresh checkout for the new session (unnecessary `bun install` + `.codegraph/` rebuild cost) | Don't 鈥?for _general_ launches: a session is a conversation pointed at a directory, so reuse the current workspace; worktree creation is a separate _workflow_ decision the launched session makes if/when it starts a workflow. **Ticket-recipe carve-out ([org-internal #2297]):** the ticket's work happens in a dedicated worktree anyway (worktree discipline), and claim-first pushes the workflow branch before launch 鈥?so the launcher pre-builds the worktree as the claim carrier + ready workspace, and the session reuses it (cost is front-loaded, not wasted). The worktree path goes in the claim comment + seed prompt; `?directory=` stays on the creator's workspace |
| 11 | Launching a worker and never hearing back 鈥?orchestrator polls a few times, stops, loses visibility ([org-internal #2374]) | Append the completion-report block to the seed prompt (Completion-report protocol section): worker `prompt_async`s a one-line status back into the orchestrator session on done/blocked/handoff, with issue-comment fallback |
| 12 | Headless worker calls the `question` tool 鈫?no attached client: the question deadlocks in a queue, or `auto_approve` silently auto-approves a direction-setting decision ([org-internal #2378]) | NEVER call `question` from a headless worker. Decision boundary (Completion-report protocol section): directional/irreversible decision 鈫?`status=blocked` report + one-line decision point, then STOP; the orchestrator is the sole human-decision entry point |
| 13 | Headless worker self-merges its PR / self-closes the issue once the PR looks ready, despite a seed instruction to keep it open ([org-internal #2386]; N-02 [org-internal #2367] / N-03 [org-internal #2368], commits `d1565c99` / `040f21b7`) | Self-merge / self-close / touching `main` are irreversible directional actions 鈫?blocked by the Decision boundary. PR ready 鈫?keep it OPEN, report `status=done branch=<ref>`; the orchestrator opens/admits the PR, merges and closes uniformly |
| 14 | Trusting `state=generating` as "alive" 鈥?a hung stream goes unrescued (session A, 2026-08-16: `generating` 6+ min, count frozen at 478) | Poll state AND message count twice 30鈥?0 s apart; frozen count = hung 鈫?abort + re-wake (Post-launch lifecycle section) |
| 15 | Re-sending a wake prompt to a hung session and waiting 鈥?queued prompts are never consumed while the stream is dead | Abort FIRST (`POST /session/:id/abort` 鈫?200, verify `idle`), THEN re-send; a prompt queued before the abort may still never fire 鈥?always re-deliver |
| 16 | A worker woken from restart / long idle acts on its stale worldview 鈥?re-dispatches superseded work or claims authority it no longer holds (2026-08-16: revived worker re-ran an already-published FAIL verify; ops session announced an inherited "orchestrator" takeover) | Fact-baseline injection (Post-launch lifecycle section): authoritative facts + invalidated assumptions + one concrete re-assignment. Only the CURRENT authority injects; a session that suspects IT is stale verifies identity against the durable record first (`rules/compact.md` identity clause) |
| 17 | Worker goes silent mid-task: provider quota wall killed the turn (usage limit / 429) 鈥?no completion-report, no blocked report; from the issue side indistinguishable from silent completion (2026-08-29 [org-internal #3627]: opencode-go `weekly usage limit reached` after the 13:18 Round-1 FAIL; ticket stalled until manually re-woken) | Detect via tail probe (`finish=error` + 鈿狅笍 usage-limit part, `metadata.reason` in the quota family) + `GET /provider` quota markers; re-wake on a known-good provider with the fact-baseline template 鈥?`reference/quota-wall-recovery.md` ([org-internal #3669]). Prevent: pre-dispatch provider check + the seed's 閰嶉鑷姤 clause (`reference/completion-report.md`, [org-internal #3669]) 鈥?quota exhaustion is a `status=blocked reason=quota-exhausted` report, not a silent death |
## Known limitation (separate from this workaround)
A missing `model` producing an **apparent silent stall** (no surfaced error, no
generation) is an observability / robustness gap. Current code does NOT
hard-require `model` 鈥?it fills it via a fallback chain (`prompt.ts:858`) and any
resolution failure surfaces as a published `Session.Event.Error` on the
synchronous path (`prompt.ts:804-822`), but the **async** `prompt_async` handler
catches that cause (`forkPromptDelivery` in `handlers/session.ts`) so the HTTP
caller never sees it. The proper fix is to surface the error (or reject unusable
models at the boundary) so a stalled session is diagnosable. Until then, **always
provide an explicit known-good `model`** 鈥?this skill is the documented
workaround (see [org-internal #1695] "閬楃暀 minor").
**Update ([org-internal #2912], 2026-08-20)**: usage-limit turn failures (HTTP 429 /
FreeUsageLimitError / GoUsageLimitError) are no longer silent shells. When such
a turn terminates (retry schedule ends or the spin is aborted), the errored
assistant message persists `finish=error` plus a human-readable 鈿狅笍 text part
carrying provider, limit reason, and reset hint (`metadata: { error: true,
reason, provider }`) 鈥?visible via `GET /session/:id/message` and rendered in
the web UI like any assistant text. Pitfall #14's double-poll is still the
liveness check while a 429 retry spin is IN progress (the spin itself retries
with backoff until interrupted); the explicit error part only appears once the
turn ends. The general async-path observability gap above remains tracked by
[org-internal #1695].
## See also
- Source issue / full manual: **[org-internal #1695]** (this skill adds the ticket-driven recipe
- current-model reuse, and corrects the endpoint/claim against current code).
- Completion-report protocol: **[org-internal #2374]** (worker-side active report back to the
launching orchestrator session; fallback to issue comment).
- Quota-wall recovery recipe: **[org-internal #3669]** (evidence instance [org-internal #3627]) 鈥?detect a
worker killed by a provider usage limit, re-wake on a known-good provider,
pre-dispatch prevention.
- Decision boundary: **[org-internal #2378]** (headless worker MUST NOT call `question`;
directional decisions go through `status=blocked` reports to the
orchestrator) and **[org-internal #2386]** (the worker MUST NOT self-merge its PR,
self-close the issue, or touch `main` 鈥?the orchestrator merges and closes
uniformly).
- Correction of the prior misdiagnosis: **[org-internal #1691]** (closed Invalid), comment 9205.
- Synchronous client always sends `model`:
`packages/app/src/components/prompt-input/submit.ts:158-165`.
- Code citations: `<harness-package>/src/session/prompt.ts:858,804-822,1399`,
`handlers/session.ts` `forkPromptDelivery` (async failure handling),
`server/routes/instance/httpapi/groups/session.ts`.
## References
**On-demand references** (NOT injected 鈥?read at the declared timing):
- `reference/ticket-recipe.md` 鈥?ticket-driven launch (`Launch a session for a
specific ticket (#N)`): read BEFORE creating any session for issue #N 鈥?
Step-0 owner check, seed-prompt composition, launch + poll commands.
- `reference/completion-report.md` 鈥?completion-report protocol (瀹屾垚鍥炴姤鍗忚,
[org-internal #2374]): read BEFORE composing a seed prompt for a worker you need to hear
back from 鈥?orchestrator block template, worker-side report steps, decision
boundary ([org-internal #2378] / [org-internal #2386]), quota self-report clause ([org-internal #3669]).
- `reference/lifecycle-ops.md` 鈥?post-launch lifecycle ([org-internal #2459]): read when
managing a launched session 鈥?health-check polling, hung-stream diagnosis
(state=generating + frozen counters), abort + re-wake rescue,
fact-baseline injection, authorization asymmetry.
- `reference/quota-wall-recovery.md` 鈥?quota-wall recovery ([org-internal #3669], evidence
[org-internal #3627]): read when a worker goes silent mid-task (suspected provider usage
limit) or BEFORE dispatching a long-running worker 鈥?detection signals
(finish=error 鈿狅笍 tail, provider quota markers, log grep patterns),
provider-switch re-wake + fact-baseline template, pre-dispatch provider
check, quota self-report clause.
- `rules/headless-session-ops` wiki page 鈥?authoritative reference (this
skill is its runtime carrier; link in the Directory-convention note above).
- `core/rules/session-scope-guard.md` 鈥?owner-check + claim-first rule
(the ticket recipe's Step 0 is its launch-time application).
- `core/rules/compact.md` 鈥?identity-verification clause backing the
authorization-asymmetry rule (`reference/lifecycle-ops.md`).
@@ -0,0 +1,173 @@
> Extracted from headless-session-ops/SKILL.md (Completion-report protocol (完成回报协议, [org-internal #2374])) — moved verbatim 2026-08-26, ticket [org-internal #3480].
## Completion-report protocol (完成回报协议, [org-internal #2374])
`prompt_async` is fire-and-forget **for the launcher too**: once the worker
session starts generating, the orchestrator receives NO event when it finishes
or blocks. The `task` tool auto-returns `task_result` for subagents; a
`prompt_async`-launched independent main session has no equivalent — the
orchestrator's only built-in channel is bare polling of
`GET /session/:id/message?limit=1` (tail probe — the no-`limit` form loads the
whole transcript server-side), which it eventually stops doing ([org-internal #2366]: worker ran
~1h past the orchestrator's last poll, merged PR [org-internal #2371] and closed [org-internal #2366] while
the orchestrator had zero visibility).
The protocol below closes that gap WITHOUT changing server semantics: the
**worker** actively posts a completion/block message back into the
orchestrator's session via `prompt_async`. Fire-and-forget launch is
unchanged; the report is a worker-side obligation, not a new server dependency.
### Orchestrator side — pass your SID in the seed prompt
When launching a worker that you need to hear back from (DAG node execution,
parallel chunk work, any long-running delegated ticket), append a
completion-report block to the seed prompt. It MUST state:
- the orchestrator's **session id** (`$LAUNCHER_SID`) and **directory**;
- the report endpoint: `POST $BASE/session/$LAUNCHER_SID/prompt_async?directory=$DIR` (or, when the directory is unknown to the worker, the global route `POST $BASE/prompt_async` with the sessionID in the body — [org-internal #4307]);
- the exact report format (one-line, machine-greppable);
- the fallback (issue comment) if the report call fails.
Template (include verbatim in `seed.txt` when you are the launcher):
```text
## 完成回报协议 (mandatory, [org-internal #2374])
你由编排会话 (orchestrator) 启动。完成、阻塞、或将工单移交他人时,必须向编排会话回报:
POST <BASE>/session/<LAUNCHER_SID>/prompt_async?directory=<DIR>
body: {"agent":"builder","model":{...},"parts":[{"type":"text","text":"<report>"}]}
<report> 格式(单行,machine-greppable):
[#<ticket>] [<node-or-slug>] status=done|blocked|handoff branch=<ref> [pr=#<N>] verify=<changed+typecheck> risk=<high|low> summary=<one line>
默认(TD-678 / [org-internal #4425]):worker 只推分支不开 PR——交付推送前置本地验证
<harness-package> 下 test:changed 全绿 + typecheck 0 errorverify 字段回执);
risk 按 risk-classifier 冻结表(HIGH_RISK_GLOBS)自报。PR 由编排按容量串行开
1:1(正文携 worker 报告),或 batch 模式 epic 经 land-batch 技能按迭代统一合成
(见 core/skills/land-batch/)。pr=#<N> 仅限编排不可达自开(uncoordinated)。
若回报调用失败(HTTP 非 2xx / 连接拒绝 — orchestrator 会话可能已结束或实例重启),
退而在源 issue(或父 Epic)发同内容评论作为兜底(fallback, 非主通道)。
回报是你的 mandatory step,与 issue-checklist-sync 并列;不得静默结束。
决策边界([org-internal #2378]):headless worker 禁止调用 `question` 工具 —— 遇方向性/不可逆
决策点 → `status=blocked` + 一行决策点描述后 STOP,勿调 question(无人应答死锁 /
被 auto_approve 静默自动批准)。
配额自报([org-internal #3669]):若收到配额/用量上限错误(usage limit / 429 / 402
insufficient_quota),立即回报 status=blocked reason=quota-exhausted
summary=<provider 与 ⚠️ 部分中的重置提示>branch=<ref> 必带)后停止
——勿自行换模型重试;换供应商重唤醒由编排者执行
reference/quota-wall-recovery.md)。
PR 就绪后保持 open:不自行合并 PR、不关闭 issue、不动 main[org-internal #2386]);
合并与 issue 关闭由编排者统一执行。
```
The orchestrator discovers its own `$LAUNCHER_SID` from context (it is the
session the orchestrator IS — e.g. the SID it already used for the
`GET /session/$LAUNCHER_SID` model lookup in the recipe above) or from its
session record.
### Worker side — report on completion / block / handoff
A session whose seed prompt contains the completion-report block MUST, as its
final step (before idling):
1. **Report** by `prompt_async` into the orchestrator session, with the
one-line format above. Include `branch=<ref>` (plus `verify=`/`risk=` on
done); `pr=#N` only for an `uncoordinated` self-open; use
`status=blocked` + the blocker description when stuck; use
`status=handoff` when deliberately transferring the ticket.
1. **Quota exhaustion is a blocked report, not a silent death ([org-internal #3669]).** A
worker whose turn is killed by a usage-limit error (429 / 402 quota
family) cannot finish the turn — but per the seed's 配额自报 clause it
reports `status=blocked reason=quota-exhausted summary=<provider +
reset hint>` (with `branch=<ref>`) BEFORE stopping
(issue-comment fallback if the orchestrator session is unreachable).
The orchestrator then runs the provider-switch recovery in
`reference/quota-wall-recovery.md` instead of discovering the wall by
polling.
2. **On failure** of that POST (non-2xx / connection refused — the
orchestrator may have ended or the instance restarted), fall back to a
`工单评论 API(见 TERMINOLOGY` on the source issue / parent Epic carrying
the same one-line report. The issue comment is the durable record of last
resort, NOT the primary channel — the orchestrator does not watch issues in
real time.
3. The report is **additive** to existing obligations (issue checklist sync,
`## 当前状态` updates, archive-at-close) — it never replaces them. The issue
remains the stakeholder-facing record; the report is the
orchestrator-facing wake-up signal.
### Orchestrator side — serial PR admission on done reports ([org-internal #4425])
On a `status=done branch=<ref> verify=… risk=…` report:
1. Enqueue the ticket; admit PRs serially per `ticket-lifecycle.md` §PR
准入 — one open PR at a time, next only after double-green merge. Branch
rot is bounded by the existing keep-mergeable/syncMain machinery, not new
code; keep the queue shallow.
2. Open the 1:1 PR (title from the ticket node, body carrying the worker
report: branch / files / self-test / verify), apply the `Risk/*` label
from the report's `risk=` hint, merge via `script/pr-merge.sh` ([org-internal #3864]).
Batch-mode epics: one `land-batch` PR per iteration instead ([org-internal #3731]).
3. A first red on the opened PR goes back through the report fallback
channel (issue comment) — the worker may already be idle; the issue
comment is the durable re-entry point.
Prose discipline on the orchestrator session only — no new daemon, no
mergeable-state automation (merge-coordinator retired, [org-internal #4385]).
### Decision boundary (决策边界, [org-internal #2378])
The report channel defines how the worker *answers*; this subsection defines
the boundary of what it may *ask*. A headless worker session has NO attached
client (no Web UI / TUI), so calling the `question` tool has exactly two
outcomes, both structurally broken:
- the question sits in a queue nobody answers → the session deadlocks;
- or it is silently auto-approved by `auto_approve` → a direction-setting
decision gets decided by a default value, violating the human-gate
semantics of the review gates.
**Rule: a headless worker MUST NOT call the `question` tool.** Decision
authority is layered instead:
- **Directional / irreversible decisions** (scope changes, mid-stream
reclassification — big-bug relabel / DAG re-derivation, merge timing,
contract changes, approach selection) → the
worker reports `status=blocked` with a one-line description of the decision
point, then STOPs. The orchestrator is the sole human-decision entry point:
it interacts with the user, then replies to the worker (or acts on its
behalf).
- **No self-merge / no self-close / never touch main ([org-internal #2386])**: a headless
worker MUST NOT merge its own PR, close the source issue, or rebase /
force-push `main`. These are irreversible directional actions of exactly
the class this boundary blocks — even when the seed prompt's instruction
to "keep the PR open" is absent or the worker judges the PR ready. When
the PR is ready the worker keeps it OPEN and reports
`status=done pr=#N` to the orchestrator; the orchestrator performs the
merge and the issue close uniformly. (Counter-example: N-02 [org-internal #2367] /
N-03 [org-internal #2368] self-merged their PRs — commits `d1565c99` / `040f21b7`
despite a seed instruction to keep them open.)
- **Local implementation decisions within the AC scope** (pattern choice,
helper extraction, test shape) → the worker decides autonomously and notes
the choice in its report / PR body.
Rule of thumb: the worker's only way to "ask" is a `status=blocked` report.
A genuine ambiguity left by the seed prompt that would change scope is a
blocked report, NOT a `question` call. (Field sample: N-04 [org-internal #2369] — the
worker spontaneously did NOT merge its own PR and left merging to the
orchestrator; this subsection codifies that behavior.)
### Semantics preserved
- **Fire-and-forget launch is unchanged**: Step 13 above are identical; the
orchestrator still gets no server push at launch time.
- **No server change**: the report reuses the existing `prompt_async` endpoint
against the orchestrator's SID. Nothing new is required of the backend.
- **Polling stays as fallback**: an orchestrator that never got a report can
still poll `/session/:id/message`; the protocol removes the *need* to poll
forever, not the ability.
### (Optional / future) server-side callback
Longer-term, `prompt_async` could accept a `callback_session_id` so the server
itself posts an event to the callback session when the run loop terminates —
removing the prompt-level manual protocol. That is a separate platform
enhancement and intentionally NOT in this ticket's scope; the manual protocol
above works today and remains compatible with a future automatic callback.
@@ -0,0 +1,106 @@
> Extracted from headless-session-ops/SKILL.md (Post-launch lifecycle: health check, hung rescue, stale-worldview correction ([org-internal #2459])) — moved verbatim 2026-08-26, ticket [org-internal #3480].
## Post-launch lifecycle: health check, hung rescue, stale-worldview correction ([org-internal #2459])
Launching is half the job — a headless worker can hang silently or wake up
with an outdated worldview. All three failure modes below were observed in
production on 2026-08-16 during Epic [org-internal #2459] (ops-restart window); the runbooks
are extracted from that incident record.
### Health check — three states, one poll pattern
Poll two signals together — `GET /session?directory=` for `.state` (plus the
per-session `progress` object when present), and `GET
/session/:id/message?limit=1` for the newest message's `.info.id` — **twice,
spaced 3060 s**. `MessageID` is monotonic (`MessageID.ascending`,
`message.ts`), so a frozen newest-id across both samples is exactly the old
"message count frozen" signal at O(1) per poll instead of O(transcript) —
never poll the no-`limit` form for liveness, it loads the whole transcript
server-side:
| Verdict | Signals | Action |
|---|---|---|
| healthy | `state=generating` AND (`progress.stepCount` growing OR newest message id advancing) | leave it alone |
| idle | `state=idle` AND `progress` absent-or-stale across both polls | turn ended — read the LAST message (same `limit=1` fetch): task unfinished → wake prompt; `status=done` report → harvest |
| **hung** | `state=generating` AND `progress.stepCount`/`lastStepAt` AND newest message id ALL frozen across both polls | the generation stream is dead; queued prompts will NEVER be consumed — rescue required |
`state` alone lies in BOTH directions ([org-internal #3215]): a hung session still reports
`generating` (field sample: session A, 2026-08-16 — `state=generating` for 6+
minutes with the message count frozen at 478), and a healthy mid-step worker
can read `idle`/empty message tails in a single snapshot (2026-08-23 W3 wave:
4 healthy workers aborted off one snapshot). **The dual-sample delta is the
discriminator; a single snapshot is NEVER an abort basis.**
`progress` (`stepCount` cumulative LLM-round counter, `lastStepAt` epoch-ms
heartbeat — [org-internal #3215]) moves on every round even when message tails are
transiently empty or `state` flickers; it is absent for sessions that never
ran since instance start (treat absent = no signal, fall back to count
deltas). `lastStepAt` freshness alone does NOT prove liveness (a long tool
call inside one round keeps it stale for minutes) — always compare TWO
samples spaced ≥30 s.
### Hung rescue — abort, then re-wake (in this order)
```sh
# 1. Abort the dead stream (queued-but-unconsumed prompts do NOT unblock it)
curl -s -X POST "$BASE/session/$SID/abort?directory=$DIR" -o /dev/null -w "%{http_code}\n" # → 200
# 2. Verify idle
curl -s "$BASE/session?directory=$DIR" | jq -r --arg s "$SID" '.[] | select(.id==$s) | .state' # → idle
# 3. Check the worktree — the hung turn may have left uncommitted files (NOT lost)
git -C <worktree> status --short
# 4. Re-deliver the wake prompt (same Iron Law: known-good model).
# A prompt queued BEFORE the abort may still never fire — always re-send.
```
Endpoint: `POST /session/:sessionID/abort` (`groups/session.ts:113`,
`handlers/session.ts:311`). Abort stops the run loop; filesystem writes the
hung turn already made survive — inspect the worktree and list any recovered
files IN the wake prompt so the worker re-validates them instead of redoing
work (field sample: session A's hung turn had produced 3 src + 1 test file
that its revived self adopted).
**Anti-re-hang clause — include in every wake prompt.** The trigger for the
observed hang was a system-injected "请在适当的时机压缩当前会话" (compact at an
appropriate time): the worker ended its turn after compacting, leaving the
task half-done and idle. A wake prompt MUST carry, verbatim:
```text
若系统再注入「请在适当的时机压缩当前会话」:执行压缩后立即在后续 turn 继续任务,
绝不在任务未完成时以 idle 结束。
```
### Stale worldview — fact-baseline injection (correct BEFORE it acts)
A worker woken after an ops restart / long idle carries the worldview it went
to sleep with. It may re-dispatch superseded work, overwrite newer state, or
claim authority it does not hold (field samples, 2026-08-16: a revived worker
re-ran an M-01 verify another session had already published as FAIL —
deduplicated via flag #comment-21590; an ops-notification session inherited an
"orchestrator" identity from a compaction summary and announced a takeover —
corrected via 勘误 #comment-21654/[org-internal #21668]).
When you detect a stale-worldview session, do not wait for it to finish being
wrong — inject a fact-baseline prompt immediately (regular `prompt_async`),
structured as:
1. **You were woken; your worldview is stale** — name the event (restart /
maintenance window) and the current time.
2. **Authoritative state** — numbered facts with artifact links (wiki page,
issue comment), each with its timestamp; state explicitly which of the
recipient's standing assumptions are now INVALID.
3. **Your actual assignment now** — one concrete task (or explicit standby).
4. **Evidence rule** — verify each fact at its cited source before acting;
never act on this baseline alone.
(Field sample: the 2026-08-16 injection to session B pivoted it from the
superseded M-01 verify to the N-04b fix within one turn — the format works.)
### Authorization asymmetry — read this BEFORE "correcting" anyone
Only the session that CURRENTLY holds the authority may inject a baseline or
re-task a worker. If YOU might be the stale one — you woke from a restart,
your context came from a compaction summary, you cannot find your claim in
the durable record — assume YOU are stale: verify your identity/authority
against the record (issue assignee, claim comment, orchestrator session id)
BEFORE issuing any instruction. See the identity-verification clause in
`core/rules/compact.md` (recovery contract).
@@ -0,0 +1,226 @@
> Added from headless-session-ops/SKILL.md (triggers `quota wall` / `配额墙`,
> Post-launch lifecycle section, pitfall #17, References) — new runbook
> 2026-08-29, ticket [org-internal #3669] (evidence instance [org-internal #3627]). Companion runbooks:
> `lifecycle-ops.md` (hung rescue, fact-baseline structure) and
> `completion-report.md` §"Orchestrator side" (the quota self-report clause);
> §Prevention is this file's own closing section.
## Quota wall — worker killed mid-task by a provider usage limit
A **quota wall** is when a headless worker's turn is killed by a terminal
provider-quota error (usage limit / 429 family) mid-task. From the issue side
it looks EXACTLY like silent completion: no completion-report, no
`status=blocked` report, the ticket just stops moving. The stalled window
equals the orchestrator's polling interval, and without this runbook the
recovery was improvised.
**Evidence instance ([org-internal #3627], 2026-08-29)** — the shapes below are field-verified:
|时刻 (+08)| 事件 |
|---|---|
| 08-29 13:18 | review-code Round-1 synthesis (FAIL) posted on [org-internal #3627]; worker session `ses_fb46e705bffe9kYbSoHayoJ1OQ` (builder, opencode-go model, branch `workflow/session/3627-wedge-promptasync`) entered its revision loop |
| ~13:1816:00 | worker hit the provider's `weekly usage limit reached` mid-revision — turn killed, **no** completion-report, **no** blocked report; issue side silent |
| detection | orchestrator noticed the `/session` state stall, then confirmed via the quota error in the session log (message tail carries the ⚠️ usage-limit part, [org-internal #2912]) |
| recovery | re-wake via `prompt_async` on a known-good provider (`zhipuai-coding-plan`/glm-5.3 — since renamed `zai-coding-plan`) **plus a fact-baseline injection** (what happened / what is done / where to resume) |
| 16:26 / 16:51 | Round-2 synthesis PASS; `status=done pr=[org-internal #3675]` report received |
## Detection — three signals, in this order
Set the usual variables first:
```sh
BASE=http://127.0.0.1:4096 # dev backend (:4180 prod)
DIR=/data/octopus
SID=<worker session id>
```
**Signal 1 — tail probe (primary; survives restarts, lives in the DB).** Since
[org-internal #2912]/[org-internal #3190] a quota-killed turn persists an assistant message with
`finish=error` plus a human-readable ⚠️ text part whose metadata names the
quota family. Probe the TAIL (`limit=1`, never the no-`limit` full transcript):
```sh
curl -s "$BASE/session/$SID/message?directory=$DIR&limit=1" | jq '
{role: .[0].info.role, finish: .[0].info.finish,
error_parts: [.[0].parts[]? | select(.type=="text" and .metadata?.error == true)
| {text: .text[0:160], metadata: .metadata}]}'
```
Quota-wall verdict: `role=assistant` AND `finish=error` AND an error part with
`metadata.reason` in `{account_rate_limit, free_tier_limit, account_usage_limit,
rate_limit}` (metadata also carries `provider`). The ⚠️ text is one of — all
strings observed in this repo's runtime (`session/retry.ts`):
- `` `weekly usage limit` reached. It will reset in … `` / `Go limit reached`
(GoUsageLimitError — the [org-internal #3627] shape; `weekly` is the body's `limitName`)
- `Free limit reached` (FreeUsageLimitError)
- `Account-level usage limit — switching API keys does not reset it. You can
continue after HH:mm.` ([org-internal #3407], zhipu/zai code 1308 — account-level)
- `Provider <id> rate limited (HTTP 429)` (generic 429)
- `Generation failed — …` with a 402 `insufficient_quota` body (balance, not
window — recovery is top-up, not provider switch)
**Signal 2 — state (two shapes; use the lifecycle-ops dual-sample rule).**
```sh
curl -s "$BASE/session?directory=$DIR" | jq -r --arg s "$SID" \
'.[] | select(.id==$s) | .state'
```
- **idle** + Signal-1 tail → the turn already died on the quota error. No
abort needed; go straight to Recovery.
- **generating** with `progress`/newest-message-id frozen across two polls
3060 s apart → the 429 retry spin is still running (it retries with backoff
until the schedule ends) OR the stream is hung — treat exactly like the hung
shape in `lifecycle-ops.md`: abort FIRST, then re-wake. A single snapshot is
never an abort basis ([org-internal #3215]).
**Signal 3 — provider-level confirmation ([org-internal #2911] quota markers).** `GET
/provider` overlays active usage-limit cooldowns, so the dead provider is
visible without reading logs:
```sh
curl -s "$BASE/provider?directory=$DIR" | jq '[.all[] | select(.quota) | {id, quota}]'
# quota: {markedAt: <epoch-ms>, resetAt?: <epoch-ms>} — resetAt absent = reset unknown
date -d @$(($(curl -s "$BASE/provider?directory=$DIR" | jq '[.all[].quota.resetAt // 0] | max') / 1000)) # earliest full-reset wall clock
```
**Log grep (fallback; the dev log is restart-truncated).** Patterns actually
seen in the wild — [org-internal #3627] produced the first:
```sh
grep -E "usage limit reached|Go limit reached|Free limit reached|Account-level usage limit|GoUsageLimitError|FreeUsageLimitError|rate limited \(HTTP 429\)|insufficient_quota" \
/tmp/octopus-backend.log | tail -20
# the pool-failover WARN precedes the wall when same-name pools exist:
# "Usage limit reached on <provider> — retrying <model> on <provider>"
```
The wall is only terminal after failover has exhausted the same-name pools
([org-internal #2911] failover); the `retrying … on …` WARNs above tell you the pool was
already draining before it died.
## Recovery — abort (only if needed), re-wake on a known-good provider, inject facts
**Step 0 — classify the state (Signal 2).** `idle` → skip abort.
`generating` + frozen across dual samples → abort first, exactly per
`lifecycle-ops.md` (queued prompts are never consumed by a dead stream):
```sh
curl -s -X POST "$BASE/session/$SID/abort?directory=$DIR" -o /dev/null -w "%{http_code}\n" # → 200
curl -s "$BASE/session?directory=$DIR" | jq -r --arg s "$SID" '.[] | select(.id==$s) | .state' # → idle
```
**Step 1 — inventory surviving work.** The killed turn's filesystem writes
survive; list them so the worker re-validates instead of redoing:
```sh
git -C <worktree> status --short
git -C <worktree> log --oneline -3
```
**Step 2 — pick a known-good provider/model.** The re-wake payload MUST carry
an explicit `model` (SKILL.md Iron Law), and after a quota wall the CURRENT
provider is by definition NOT known-good. Pick a provider that is connected
AND not quota-marked:
```sh
curl -s "$BASE/provider?directory=$DIR" | jq -r \
'.connected as $c | [.all[] | select(.id as $id | $c | index($id)) | select(.quota | not) | .id] | join("\n")'
# then pick a modelID you know works on that provider (e.g. the orchestrator's
# own current model — it is generating by definition).
```
([org-internal #3627]: opencode-go dead → re-woken on zhipuai-coding-plan/glm-5.3, since
renamed `zai-coding-plan`. Do NOT "wait for reset" as the default — weekly
windows mean days.)
**Empty list — every connected provider is quota-marked.** The loop needs an
explicit exit; never re-wake into a wall you can see. Find the earliest reset
among the marked connected providers:
```sh
curl -s "$BASE/provider?directory=$DIR" | jq -r \
'.connected as $c
| [.all[] | select(.id as $id | $c | index($id)) | select(.quota)]
| min_by(.quota.resetAt // 9007199254740991) # resetAt absent → sorts last
| "\(.id): markedAt=\(.quota.markedAt) resetAt=\(.quota.resetAt // "unknown")"'
# resetAt is epoch-ms → wall clock: date -d @$((<resetAt> / 1000))
```
- Known `resetAt` within reach → suspend the re-wake (leave the worker down)
and re-run this Step after the reset; weekly windows mean days — schedule
the retry, do not poll in a tight loop.
- `resetAt` unknown, or the earliest window is unacceptable (402
`insufficient_quota` is balance — reset means top-up, not time) →
**escalate to the human orchestrator**: report every marked provider's
`id`/`markedAt`/`resetAt` and stop. Top-up or onboarding a new provider is
a human decision, not another re-wake.
**Step 3 — re-wake with a fact-baseline prompt.** Build the payload from a file
(`jq --rawfile`, SKILL.md pitfall #7), same wake shape as
`lifecycle-ops.md` but with the quota facts. The template also restates the
completion-report protocol ([org-internal #2374]) so the recovery is self-contained: it
assumes the original seed still contains the protocol block — if you are
unsure it does (e.g. the worker was dispatched before [org-internal #3669] added the quota
clause), the template's own block below carries it; fill
`<BASE>`/`<LAUNCHER_SID>`/`<DIR>` with the orchestrator session's values:
```sh
cat > wake.txt <<'EOF'
## 事实基线(配额墙恢复)
你上一回合被供应商配额错误杀死,世界观停留在被杀时刻。先读以下事实再行动:
1. 发生了什么:<HH:mm> 你的回合因 <provider> 配额耗尽终止(usage limit
预计重置 <HH:mm 或 未知>);本消息由已知良好供应商重新唤醒,非你自行恢复。
2. 已做什么:<已完成步骤/commit/PR 状态>worktree 未提交变更:
<git status --short 摘要>——先 re-validate 这些文件,勿重做已完成工作。
3. 从哪继续:<一件具体的下一步任务>。
4. 证据规则:逐条到引用来源(issue 评论 / PR / 工件)核实后再行动,
勿仅凭本基线行动。
## 完成回报协议(重申 [org-internal #2374]/[org-internal #3669];不确定原始 seed 是否含协议块时以本节为准)
完成、阻塞或移交时,必须回报编排会话(你的唤醒者):
POST <BASE>/session/<LAUNCHER_SID>/prompt_async?directory=<DIR>
body: {"agent":"builder","model":{...},"parts":[{"type":"text","text":"<report>"}]}
<report> 单行格式:[#<ticket>] [<node-or-slug>] status=done|blocked|handoff pr=#<N> summary=<一行>
约束:若再遇配额/用量上限错误,立即回报 status=blocked
reason=quota-exhausted summary=<provider 与重置提示>pr=#<N> 有 PR 时带上)
后停止,勿自行换模型重试;回报失败 → 源 issue 评论兜底。
EOF
jq -n --rawfile p wake.txt '{agent:"builder",
model:{providerID:"<known-good providerID>",modelID:"<known-good modelID>"},
parts:[{type:"text",text:$p}]}' \
| curl -s -X POST "$BASE/session/$SID/prompt_async?directory=$DIR" \
-H 'content-type: application/json' -d @- -o /dev/null -w "%{http_code}\n" # → 204
```
Then poll the tail (`limit=1`) until an assistant message appears, and keep
expecting the completion-report per protocol — the re-woken worker owes you
`status=done|blocked` like any other (the restated block in the template
makes this hold even when the original seed predates [org-internal #3669]).
## Prevention — before dispatch, and in the seed
**Pre-dispatch provider check (30 seconds, catches most walls):**
```sh
# 1) no active quota markers on the dispatch provider (else pick another)
curl -s "$BASE/provider?directory=$DIR" | jq '[.all[] | select(.quota) | .id]'
# 2) dispatch provider is connected AND ≥1 other connected provider exists as fallback
curl -s "$BASE/provider?directory=$DIR" | jq '{connected, fallbacks: (.connected | length > 1)}'
```
If the intended provider is already quota-marked, dispatch on a different one
— do not launch into a wall you can see coming.
**Quota self-report clause (turns a silent death into a harvestable
blocked).** The canonical clause text lives in `completion-report.md`
§"Orchestrator side" (配额自报, [org-internal #3669]) — append it VERBATIM to the seed
prompt's completion-report block. Single source of truth: do not fork or
restate the clause here; its blocked payload already follows the one-line
key=value spec (`status=blocked reason=quota-exhausted summary=<provider +
reset hint> pr=#N`, completion-report.md `<report>` format).
With the clause in the seed, a worker that hits the wall reports `blocked`
instead of dying silently — the orchestrator harvests the report and runs the
Recovery section directly. The clause is also why the re-wake prompt's
constraint (Step 3) repeats it: the re-woken worker must know the rule still
holds on the new provider.
@@ -0,0 +1,182 @@
> Extracted from headless-session-ops/SKILL.md (Launch a session for a specific ticket (#N)) — moved verbatim 2026-08-26, ticket [org-internal #3480].
## Launch a session for a specific ticket (#N)
The common case: an agent (or cron/CI) needs to spin up a fresh main session to
work a tracked issue. Compose the seed prompt **from the issue itself** so the
new session starts with real context, and launch it on **the current session's
model**.
0. **Owner check (MANDATORY pre-step, [org-internal #1803]).** Before creating anything,
verify no other live session already owns this ticket or code area —
duplicate ownership is how [org-internal #1744]/[org-internal #1753] collided. The check spans four
data sources; any live claim → **ABORT the launch and report the
conflict**.
**Mechanized ([org-internal #3667])**: `bash script/session-conflicts.sh <N>` runs the
session-title scan (a), the branch scan (d), and the open-PR check (c)
in one read-only pass (exit 4 = conflict, `--json` for a machine
summary); the assignee/claim check (b) is `claim.sh`'s compare-and-swap
at claim time. The manual recipes below remain the fallback and the
normative definition (`core/rules/session-scope-guard.md`).
**a. Session-title scan (live session check).** The listing endpoint
returns at most `limit` (default 100) sessions — on busy repos (300+ live
sessions in one directory) that silently truncates and the scan misses
owners ([org-internal #3190]). Non-range listing responses carry `X-Total-Count` (full
filter population, untruncated) and `X-Has-More: true` (only when the
page is truncated). Scan procedure: fetch page 1, and when
`X-Has-More: true` is present re-fetch once with `limit=$TOTAL` — treat
any failure to obtain the full population as "more owners may exist"
(ABORT or widen the scan — never assume the first page is the whole
population).
```sh
# N must be the digits-only issue number (e.g. N=1803) — never interpolate
# raw issue text here; the regex below assumes digits.
[[ "$N" =~ ^[0-9]+$ ]] || { echo "N must be digits only"; exit 1; }
HDR=$(mktemp)
BATCH=$(curl -sD "$HDR" "$BASE/session?directory=$DIR")
TOTAL=$(awk -F': ' 'tolower($1)=="x-total-count"{print $2}' "$HDR" | tr -d '\r')
if [[ "$TOTAL" =~ ^[0-9]+$ ]] && awk -F': ' 'tolower($1)=="x-has-more"{print $2}' "$HDR" | grep -q true; then
# truncated page + known population — refetch the full set in one request
BATCH=$(curl -s "$BASE/session?limit=$TOTAL&directory=$DIR")
elif awk -F': ' 'tolower($1)=="x-has-more"{print $2}' "$HDR" | grep -q true; then
# truncated page but no usable X-Total-Count (older server) — the
# population is unknown; do NOT re-fetch with the same default limit
# (it would silently re-truncate). Escalate instead.
echo "ERROR: session listing truncated but X-Total-Count unavailable — cannot establish full owner population" >&2
rm -f "$HDR"; exit 1
fi
rm -f "$HDR"
echo "scanned $(echo "$BATCH" | jq 'length') of ${TOTAL:-?} sessions in $DIR" >&2
echo "$BATCH" | jq -r --arg n "$N" '.[] | select(.title | test("#" + $n + "([^0-9]|$)")) | "\(.id)\t\(.title)\t\(.time.updated)"'
```
- Any hit whose `time.updated` is recent (session still active) → **ABORT
the launch and report the conflict** (issue comment naming the owning
session id). Do not launch a second session for the same ticket.
**b. Assignee / claim check ([org-internal #2297]).** Read the issue via
`工单 API(见 TERMINOLOGYget(owner, repo, index: N)` (or `GET /api/v1/repos/$OWNER/$REPO/issues/$N`):
- If `assignee` is set and is **not** the launching agent → **ABORT** and
report (the ticket is already claimed).
- If a claim comment names a **different branch / session id** → **ABORT**
and coordinate on the issue before proceeding.
**c. Open-PR check.** List open PRs referencing `#N` —
`工单 API(见 TERMINOLOGYsearch(q: "#N", type: "pulls", state: "open")` (API source)
or `gitea_pull__list` — and **ABORT** if an open PR already covers the ticket.
**d. Remote branch check ([org-internal #2297]).** A local branch is invisible to other
sessions — check the remote too:
`git ls-remote origin 'workflow/*'` and scan for a branch tail covering `#N`.
Also check the repo side locally: `git worktree list` +
`git branch --list 'workflow/*'` for a branch / worktree already covering #N.
- A stale hit (session idle for hours / clearly abandoned) → do NOT
hand-post a takeover comment (retired admin workaround — unauditable);
dead-claim takeover is mechanized ([org-internal #3668]): `bash script/claim.sh
takeover <ticket> <branch> --session-id <id>`. It enforces the
evidence gate (branch ABSENT on origin AND (session 404 OR comment at
least CLAIM_TAKEOVER_STALE_DAYS old)) and exits 10 when the evidence
does not hold — a valid claim is never superseded; an idle-but-alive
session is NOT provably dead, coordinate on the issue instead (exit 4
= owned/conflict, in claim.sh and session-conflicts.sh alike).
Evidence rules: `core/rules/session-scope-guard.md`
§"Dead-claim takeover" (claim.sh header is normative).
The runtime counterpart of this check is `core/rules/session-scope-guard.md`
("One task, one owner"), injected into every session's prompt.
- **Backend hard guard ([org-internal #1989]):** `POST /session` hard-rejects duplicate-ticket
session creation with HTTP 400. The backend guard is **on by default ([org-internal #2350])**;
set `enabled: false` in config to opt out. It fails open ONLY on defect
paths (DB / config errors), never as a configured disable. The front-end
soft check above is now backed by this backend hard check for defense in
depth.
- **Claim-first ([org-internal #2297]).** Claiming a ticket is one atomic 3-step action:
(1) set the issue assignee to the working account, (2) post a claim comment
naming the workflow branch and session id, (3) push the workflow branch to
remote (`git push -u origin workflow/...`). The claim is valid ONLY when all
three steps complete — an incomplete claim is NOT a claim. Do NOT launch the
session until all three steps are done (see
`core/rules/session-scope-guard.md` §"Claim-first"). If a collision HAS
already happened, follow `rules/ownership-collision-runbook` (wiki, L2 on-demand).
- **Provisioning after claim ([org-internal #3642]) is ONE command** — claim + worktree +
Tier-1 runs scaffold:
`bash script/claim-provision.sh <N> workflow/<branch>` (runs claim.sh with
exit codes passed through, attaches the session worktree to the claimed
branch via `session-worktree.sh create <slug> --branch <branch>`, then
scaffolds `<runs-root>/{slug}/` via `<harness-package>/script/runs-init.ts`
with kind-aware exemptions; prints one JSON summary; idempotent re-runs).
Pitfall #10's pre-built worktree is exactly what it produces — put the
returned worktree path in the claim comment + seed prompt.
1. **Read the issue.** Use the REST API (`gitea-rest` skill; `gitea-mcp` is retired):
```sh
curl -s "$GITEA/api/v1/repos/$OWNER/$REPO/issues/$N" | jq '{title, body}'
```
or `工单 API(见 TERMINOLOGYget(owner, repo, index: N)`.
2. **Compose the seed prompt** from the issue body. The seed MUST, at minimum:
- restate the **goal** in one sentence;
- list the **constraints** and **acceptance criteria**;
- cite every **wiki artifact path** / related issue referenced in the body
(e.g. `{epic-slug}/dag`, `{epic-slug}/dag-nodes/{node-id}`, `[org-internal #1691]`);
- tell the new agent to follow the right pipeline skill for the work type
(implement / bugfix / design / …).
Write it to a file (handles newlines):
```sh
cat > seed.txt <<'EOF'
Work issue #N: <one-line title>.
Goal: <…>.
Constraints: <…>.
Acceptance: <…>.
Artifacts: <wiki paths from the issue body>.
Follow the `implement` skill (Mode: bugfix) for this.
Session scope guard (mandatory): if a test fails and it was NOT caused by
your change — classify (pre-existing/flaky), file a BF/FT issue per
core/rules/session-scope-guard.md, and continue this task. Do NOT fix
unrelated failing tests in place.
## 完成回报协议 (mandatory, [org-internal #2374])
完成、阻塞、或移交时,向编排会话回报:
POST <BASE>/session/<LAUNCHER_SID>/prompt_async?directory=<DIR>
(编排目录未知时可用全局路由 POST <BASE>/prompt_asyncbody 携 sessionID[org-internal #4307]
body: {"agent":"builder","model":{...},"parts":[{"type":"text","text":"<report>"}]}
<report>: [#N] [<node-or-slug>] status=done|blocked|handoff branch=<ref> [pr=#<PR>] verify=<changed+typecheck> risk=<high|low> summary=<one line>
约束:worker 只推分支不开 PR(PR 由编排串行开);交付推送前 test:changed+typecheck 须绿。若再遇配额/用量上限错误,立即回报 status=blocked reason=quota-exhausted summary=<provider 与重置提示>branch=<ref> 必带)后停止,勿自行换模型重试。
回报失败(非 2xx / 连接拒绝)→ 在本 issue(或父 Epic)发同内容评论兜底。
决策边界([org-internal #2378]):遇方向性决策点 → status=blocked 回报,勿调 question。
EOF
```
3. **Determine the known-good model** (reuse the current session's — see the
"Reuse the current session's model" section).
4. **Run the 3-step flow** with `title:"#N — <short>"` and the file-built
payload. Read the model from the launcher session's own record — this also
applies the `model.id → modelID` remap from the "Reuse" section:
```sh
MODEL=$(curl -s "$BASE/session/$LAUNCHER_SID?directory=$DIR" \
| jq -c '.model | {providerID, modelID: .id}')
SID=$(curl -s -X POST "$BASE/session?directory=$DIR" \
-H 'content-type: application/json' \
-d "{\"title\":\"#$N — <short>\",\"agent\":\"builder\"}" | jq -r .id)
PAYLOAD=$(jq -n --argjson m "$MODEL" --rawfile p seed.txt \
'{agent:"builder",model:$m,parts:[{type:"text",text:$p}]}')
curl -s -X POST "$BASE/session/$SID/prompt_async?directory=$DIR" \
-H 'content-type: application/json' -d "$PAYLOAD"
```
5. **Poll** Step 3 until an `assistant` message appears — bounded tail poll
(`&limit=1`), never the no-`limit` full-transcript form. Record `$SID` somewhere
durable (issue comment, CI log) so the session is traceable to the ticket.
6. **Add the completion-report block to the seed prompt** (see next section) so
the worker reports back when it finishes or blocks — without it, this
orchestrator has no push channel and must poll forever ([org-internal #2374]).