Files
octopus-workflow/core/skills/writing-skills/SKILL.md
T

13 KiB

name, description, triggers, role
name description triggers role
writing-skills Use ONLY when creating, editing, or evaluating an octopus skill — writing the SKILL.md body, deciding whether a skill needs baseline testing, or running a blind A/B eval of a behavior-shaping technique. Use BEFORE first deployment of any skill that must hold a process gate under pressure.
blind A/B
A/B eval
skill eval
技能编写
编写技能
技能评估
评估技能
行为塑造
SKILL.md
Producer

Writing Skills

Overview

Writing a skill IS test-driven development applied to process documentation. You write test cases (blind pressure scenarios run through Explorer subagents), watch them fail (baseline behavior without the skill), write the skill, watch tests pass (agents comply), and refactor (close loopholes).

Core principle: If you didn't watch an agent fail without the skill, you don't know if the skill teaches the right thing — or teaches anything at all.

This skill covers the methodology. For file locations, frontmatter fields, and valid config shapes, load customize-octopus — a built-in skill (packages/octopus/src/skill/prompt/customize-octopus.md), not a skill under .octopus/skills/. The two co-trigger: mechanics + methodology.

The Iron Law

NO BEHAVIOR-SHAPING SKILL WITHOUT A FAILING BASELINE FIRST.

This applies to NEW skills AND EDITS that change what a skill enforces. Write the skill before testing the baseline? You no longer have a baseline — the failure you needed to document is already contaminated by knowing the fix.

No exceptions for "simple additions", "just a new section", or "documentation updates". If the edit changes agent behavior, it needs a baseline.

When to eval — and when it is overkill (decide FIRST)

Not every skill needs a blind A/B eval. Classify before authoring:

Skill type Example Eval required?
Behavior-shaping / discipline (must hold a gate under pressure) review-code (Phase C), a new "never skip X" rule YES — full blind A/B (RED-GREEN-REFACTOR below)
Technique / how-to (a method agents apply) condition-based-waiting, root-cause-tracing port Application scenarios (can the agent apply it correctly?) — baseline optional
Reference / mechanical (lookup, syntax, config shapes) config-field docs, command references NO — source-fidelity check + lint suffice
Pure port of a proven technique (verbatim, minimal localization) a 1:1 port with wording unchanged Fidelity check against source + spot-check; full re-eval only if localized substantially

The question that decides it: "Could a capable agent comply with this skill without ever reading it?" If yes (it restates convention, or the gate is already procedurally explicit elsewhere), the skill is reference-class and eval will show a ceiling — both control and treatment hold, and the eval carries no signal. Do not run a blind eval expecting to prove a reference skill works; run a fidelity check instead.

RED-GREEN-REFACTOR for skills

RED — write the failing test (baseline)

Run pressure scenarios through an Explorer subagent WITHOUT the skill (or with a control SKILL that omits only the technique under test). Document, verbatim:

  • What choices the agent made.
  • The exact rationalizations it used to skip the gate.
  • Which pressures (deadline, sunk cost, authority, exhaustion) triggered the violation.

This is "watch the test fail". You must see what agents naturally do before writing the skill, or you are writing to an imagined failure.

GREEN — write the minimal skill

Write the skill to address the specific rationalizations you observed — not hypothetical ones. Run the same scenarios WITH the skill. The agent should now comply. Do not add content for failures you did not see.

REFACTOR — close loopholes

The agent found a new rationalization in GREEN? Add an explicit counter (see Rationalization Table below). Re-test until bulletproof.

Blind A/B eval methodology

For behavior-shaping skills, the eval is a blind controlled experiment: control = SKILL without the technique; treatment = SKILL with it. The Explorer subagents running the scenarios are unaware of the condition or that they are being tested.

Artifact layout (wiki pages under _evals/{skill-name}/)

The _evals/{skill-name}/ namespace is an intentional slug-less collection namespace — eval records aggregate across skills and outlive any one pipeline run — analogous to two documented slug-convention exceptions: the audit audit/{date}/ date-slug exception (NAM 4.4, core/checklists/audit-process.md) and the retrospective _retrospectives/ exception (retrospective/SKILL.md Phase 7 — Report).

README                                   findings, conclusion, honest evidence strength,
                                         and why you stopped (or didn't)
round-1-initial/control-SKILL            baseline skill (technique omitted)
round-1-initial/treatment-SKILL          skill with the technique
round-1-initial/scenarios/scenario-<id>  blind pressure scenarios (3+ for discipline skills)
round-1-initial/cells/cell-<id>-{control,treatment}  per-cell dispatch prompts
round-1-initial/outputs/output-<id>-{control,treatment}  captured agent outputs (verbatim)
round-1-initial/rubric                   scoring rubric
round-1-initial/scoring                  scored results table
round-2-replication/*                    re-runnable half; same names, NEW scenarios

Procedure

  1. Write 3+ pressure scenarios that combine pressures (deadline + sunk cost; authority + exhaustion). Each scenario tempts the agent to skip the gate the skill enforces.
  2. Run each scenario twice, blind, via an Explorer subagent: once seeded with control-SKILL.md, once with treatment-SKILL.md. The subagent sees only its assigned SKILL + the scenario — never the other condition, never the word "test" or "experiment".
  3. Score outputs against the rubric on a 3-level verdict: HOLD (kept the gate), PARTIAL (hedged / delayed but technically held), RATIONALIZE (skipped with justification).
  4. Treatment must win the differentiating cells — cells where control and treatment split. Ceiling cells (both HOLD) carry NO signal either way: they are a positive finding (that gate is robust without the technique), not noise to eliminate by cranking pressure.
  5. Replicate: run a second round with NEW scenarios before concluding. A single round is one data point.
  6. Record honest evidence strength in the README — see calibration below.

Worked example

The canonical blind A/B run (2 rounds, 17 cells, rationalization-anticipation tables) is published at the _evals/skill-compliance-framing/ wiki namespace once its artifacts are promoted from internal storage. The experiment demonstrated a ceiling finding: the tested gate was robust in both control and treatment arms without the technique under evaluation.

Honest evidence calibration (write this in the eval README)

An eval README that claims more than the data supports is worse than no eval. State explicitly:

  • Supports: the construct the data is consistent with (e.g. "consistent with a real but construct-specific effect on redundancy-gates under pressure").
  • Does NOT prove: the limits — small differentiating N, ceiling cells, stated-decision eval vs production compliance.
  • Ceiling IS a finding: when both groups hold a gate, that gate is robust without the technique. Do not manufacture pressure to break the control — that is eval p-hacking.
  • Real arbiter: for production skills, the honest next test is often a measurement script over real iterations (e.g. script/measure-review-blockers.ts), not more simulated cells.

Stopping is a valid decision. When the ceiling pattern is clear and differentiating N stays small, further stated-decision cells have diminishing returns. Say so and stop.

Match the form to the failure

Before authoring guidance, classify the baseline failure. The form that bulletproofs one failure type measurably backfires on another.

Baseline failure Right form Wrong form
Skips/violates a rule under pressure (knows better, does it anyway) Prohibition + rationalization table + red flags Soft guidance ("prefer...", "consider...")
Complies, but output has the wrong shape (bloated prompt, buried verdict) Positive recipe: state what the output IS — its parts, in order Prohibition list ("don't restate", "never narrate")
Omits a required element from something they already produce Structural: REQUIRED field or slot in the template Prose reminders near the template
Behavior should depend on a condition Conditional keyed to an observable predicate ("if the brief exists, reference it") Unconditional rule + exemption clauses

Why prohibitions backfire on shaping problems: under a competing incentive, agents negotiate with "don't X". In head-to-head wording tests, the prohibition arm produced more of the unwanted content than the recipe arm — and trended worse than even the no-guidance control. A recipe leaves nothing to negotiate: the output matches the stated shape or it doesn't.

Rules for whichever form:

  • No nuance clauses. "Don't X unless it matters" reopens the negotiation. Express a real exception as its own conditional on an observable predicate.
  • Exemption clauses don't scope. "This limit doesn't apply to code blocks" still suppresses code blocks. Restructure so the rule can't reach the exempt part.

Rationalization table

Capture the rationalizations you observed in RED. Every excuse the baseline agents used goes in the table, paired with the reality that defeats it:

| Excuse | Reality |
|--------|---------|
| "Too simple to test" | Simple code breaks. The test takes 30 seconds. |
| "I'll test after" | Tests-after answers "what does this do?"; the baseline answers "what SHOULD it do?" |

This is the same rationalization-anticipation device already embedded in implement (including its bugfix mode) — here it is an authoring tool: you build it from your own baseline data, not from guesswork.

Micro-test wording before full scenarios

Full blind A/B runs are expensive. Verify the wording itself first with micro-tests:

  1. One fresh-context sample per call — a single-shot Explorer subagent. System prompt = the realistic context the guidance will live in; user message = a task that tempts the failure.
  2. Always include a no-guidance control. If the control doesn't exhibit the failure, there is nothing to fix — stop, don't author the guidance.
  3. 5+ reps per variant. Single samples lie.
  4. Read every flagged match manually. Template echoes and quoted counter-examples masquerade as hits; automated counts overstate both failure and success.
  5. Variance is a metric. When wording lands, reps converge on the same shape. Five different interpretations across five reps means the wording isn't binding — tighten the form before adding words.

Micro-tests verify wording; they do not replace blind A/B for discipline skills.

Authoring checklist

RED — baseline:

  • Classified the skill (behavior-shaping / technique / reference / port)
  • If behavior-shaping: wrote 3+ combined-pressure scenarios
  • Ran baseline (control SKILL) blind via Explorer subagent; documented rationalizations verbatim

GREEN — minimal skill:

  • Frontmatter name + description (description = WHEN to use, not what it does — see customize-octopus)
  • Addresses the specific baseline failures observed (not hypothetical ones)
  • Form matches the failure type (table above)
  • Ran scenarios WITH skill; agent now complies

REFACTOR — loopholes:

  • New rationalizations from testing have explicit counters
  • Rationalization table built from real baseline data
  • Re-tested until bulletproof

Eval README (if blind A/B was run):

  • States what the data supports AND does not prove
  • Records ceiling cells as a positive finding, not noise
  • Names the real arbiter (measurement script / production data) if applicable