13 KiB
name, description, triggers, role
| name | description | triggers | role | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| writing-skills | Use ONLY when creating, editing, or evaluating an octopus skill — writing the SKILL.md body, deciding whether a skill needs baseline testing, or running a blind A/B eval of a behavior-shaping technique. Use BEFORE first deployment of any skill that must hold a process gate under pressure. |
|
Producer |
Writing Skills
Overview
Writing a skill IS test-driven development applied to process documentation. You write test cases (blind pressure scenarios run through Explorer subagents), watch them fail (baseline behavior without the skill), write the skill, watch tests pass (agents comply), and refactor (close loopholes).
Core principle: If you didn't watch an agent fail without the skill, you don't know if the skill teaches the right thing — or teaches anything at all.
This skill covers the methodology. For file locations, frontmatter fields,
and valid config shapes, load customize-octopus — a built-in skill
(packages/octopus/src/skill/prompt/customize-octopus.md), not a skill under
.octopus/skills/. The two co-trigger: mechanics + methodology.
The Iron Law
NO BEHAVIOR-SHAPING SKILL WITHOUT A FAILING BASELINE FIRST.
This applies to NEW skills AND EDITS that change what a skill enforces. Write the skill before testing the baseline? You no longer have a baseline — the failure you needed to document is already contaminated by knowing the fix.
No exceptions for "simple additions", "just a new section", or "documentation updates". If the edit changes agent behavior, it needs a baseline.
When to eval — and when it is overkill (decide FIRST)
Not every skill needs a blind A/B eval. Classify before authoring:
| Skill type | Example | Eval required? |
|---|---|---|
| Behavior-shaping / discipline (must hold a gate under pressure) | review-code (Phase C), a new "never skip X" rule |
YES — full blind A/B (RED-GREEN-REFACTOR below) |
| Technique / how-to (a method agents apply) | condition-based-waiting, root-cause-tracing port |
Application scenarios (can the agent apply it correctly?) — baseline optional |
| Reference / mechanical (lookup, syntax, config shapes) | config-field docs, command references | NO — source-fidelity check + lint suffice |
| Pure port of a proven technique (verbatim, minimal localization) | a 1:1 port with wording unchanged | Fidelity check against source + spot-check; full re-eval only if localized substantially |
The question that decides it: "Could a capable agent comply with this skill without ever reading it?" If yes (it restates convention, or the gate is already procedurally explicit elsewhere), the skill is reference-class and eval will show a ceiling — both control and treatment hold, and the eval carries no signal. Do not run a blind eval expecting to prove a reference skill works; run a fidelity check instead.
RED-GREEN-REFACTOR for skills
RED — write the failing test (baseline)
Run pressure scenarios through an Explorer subagent WITHOUT the skill (or with a control SKILL that omits only the technique under test). Document, verbatim:
- What choices the agent made.
- The exact rationalizations it used to skip the gate.
- Which pressures (deadline, sunk cost, authority, exhaustion) triggered the violation.
This is "watch the test fail". You must see what agents naturally do before writing the skill, or you are writing to an imagined failure.
GREEN — write the minimal skill
Write the skill to address the specific rationalizations you observed — not hypothetical ones. Run the same scenarios WITH the skill. The agent should now comply. Do not add content for failures you did not see.
REFACTOR — close loopholes
The agent found a new rationalization in GREEN? Add an explicit counter (see Rationalization Table below). Re-test until bulletproof.
Blind A/B eval methodology
For behavior-shaping skills, the eval is a blind controlled experiment: control = SKILL without the technique; treatment = SKILL with it. The Explorer subagents running the scenarios are unaware of the condition or that they are being tested.
Artifact layout (wiki pages under _evals/{skill-name}/)
The
_evals/{skill-name}/namespace is an intentional slug-less collection namespace — eval records aggregate across skills and outlive any one pipeline run — analogous to two documented slug-convention exceptions: the auditaudit/{date}/date-slug exception (NAM 4.4,core/checklists/audit-process.md) and the retrospective_retrospectives/exception (retrospective/SKILL.mdPhase 7 — Report).
README findings, conclusion, honest evidence strength,
and why you stopped (or didn't)
round-1-initial/control-SKILL baseline skill (technique omitted)
round-1-initial/treatment-SKILL skill with the technique
round-1-initial/scenarios/scenario-<id> blind pressure scenarios (3+ for discipline skills)
round-1-initial/cells/cell-<id>-{control,treatment} per-cell dispatch prompts
round-1-initial/outputs/output-<id>-{control,treatment} captured agent outputs (verbatim)
round-1-initial/rubric scoring rubric
round-1-initial/scoring scored results table
round-2-replication/* re-runnable half; same names, NEW scenarios
Procedure
- Write 3+ pressure scenarios that combine pressures (deadline + sunk cost; authority + exhaustion). Each scenario tempts the agent to skip the gate the skill enforces.
- Run each scenario twice, blind, via an Explorer subagent: once seeded
with
control-SKILL.md, once withtreatment-SKILL.md. The subagent sees only its assigned SKILL + the scenario — never the other condition, never the word "test" or "experiment". - Score outputs against the rubric on a 3-level verdict:
HOLD(kept the gate),PARTIAL(hedged / delayed but technically held),RATIONALIZE(skipped with justification). - Treatment must win the differentiating cells — cells where control and treatment split. Ceiling cells (both HOLD) carry NO signal either way: they are a positive finding (that gate is robust without the technique), not noise to eliminate by cranking pressure.
- Replicate: run a second round with NEW scenarios before concluding. A single round is one data point.
- Record honest evidence strength in the README — see calibration below.
Worked example
The canonical blind A/B run (2 rounds, 17 cells, rationalization-anticipation
tables) is published at the _evals/skill-compliance-framing/ wiki namespace
once its artifacts are promoted from internal storage. The experiment
demonstrated a ceiling finding: the tested gate was robust in both control
and treatment arms without the technique under evaluation.
Honest evidence calibration (write this in the eval README)
An eval README that claims more than the data supports is worse than no eval. State explicitly:
- Supports: the construct the data is consistent with (e.g. "consistent with a real but construct-specific effect on redundancy-gates under pressure").
- Does NOT prove: the limits — small differentiating N, ceiling cells, stated-decision eval vs production compliance.
- Ceiling IS a finding: when both groups hold a gate, that gate is robust without the technique. Do not manufacture pressure to break the control — that is eval p-hacking.
- Real arbiter: for production skills, the honest next test is often a
measurement script over real iterations (e.g.
script/measure-review-blockers.ts), not more simulated cells.
Stopping is a valid decision. When the ceiling pattern is clear and differentiating N stays small, further stated-decision cells have diminishing returns. Say so and stop.
Match the form to the failure
Before authoring guidance, classify the baseline failure. The form that bulletproofs one failure type measurably backfires on another.
| Baseline failure | Right form | Wrong form |
|---|---|---|
| Skips/violates a rule under pressure (knows better, does it anyway) | Prohibition + rationalization table + red flags | Soft guidance ("prefer...", "consider...") |
| Complies, but output has the wrong shape (bloated prompt, buried verdict) | Positive recipe: state what the output IS — its parts, in order | Prohibition list ("don't restate", "never narrate") |
| Omits a required element from something they already produce | Structural: REQUIRED field or slot in the template | Prose reminders near the template |
| Behavior should depend on a condition | Conditional keyed to an observable predicate ("if the brief exists, reference it") | Unconditional rule + exemption clauses |
Why prohibitions backfire on shaping problems: under a competing incentive, agents negotiate with "don't X". In head-to-head wording tests, the prohibition arm produced more of the unwanted content than the recipe arm — and trended worse than even the no-guidance control. A recipe leaves nothing to negotiate: the output matches the stated shape or it doesn't.
Rules for whichever form:
- No nuance clauses. "Don't X unless it matters" reopens the negotiation. Express a real exception as its own conditional on an observable predicate.
- Exemption clauses don't scope. "This limit doesn't apply to code blocks" still suppresses code blocks. Restructure so the rule can't reach the exempt part.
Rationalization table
Capture the rationalizations you observed in RED. Every excuse the baseline agents used goes in the table, paired with the reality that defeats it:
| Excuse | Reality |
|--------|---------|
| "Too simple to test" | Simple code breaks. The test takes 30 seconds. |
| "I'll test after" | Tests-after answers "what does this do?"; the baseline answers "what SHOULD it do?" |
This is the same rationalization-anticipation device already embedded in
implement (including its bugfix mode) — here it is an authoring tool: you build it from
your own baseline data, not from guesswork.
Micro-test wording before full scenarios
Full blind A/B runs are expensive. Verify the wording itself first with micro-tests:
- One fresh-context sample per call — a single-shot Explorer subagent. System prompt = the realistic context the guidance will live in; user message = a task that tempts the failure.
- Always include a no-guidance control. If the control doesn't exhibit the failure, there is nothing to fix — stop, don't author the guidance.
- 5+ reps per variant. Single samples lie.
- Read every flagged match manually. Template echoes and quoted counter-examples masquerade as hits; automated counts overstate both failure and success.
- Variance is a metric. When wording lands, reps converge on the same shape. Five different interpretations across five reps means the wording isn't binding — tighten the form before adding words.
Micro-tests verify wording; they do not replace blind A/B for discipline skills.
Authoring checklist
RED — baseline:
- Classified the skill (behavior-shaping / technique / reference / port)
- If behavior-shaping: wrote 3+ combined-pressure scenarios
- Ran baseline (control SKILL) blind via Explorer subagent; documented rationalizations verbatim
GREEN — minimal skill:
- Frontmatter
name+description(description = WHEN to use, not what it does — seecustomize-octopus) - Addresses the specific baseline failures observed (not hypothetical ones)
- Form matches the failure type (table above)
- Ran scenarios WITH skill; agent now complies
REFACTOR — loopholes:
- New rationalizations from testing have explicit counters
- Rationalization table built from real baseline data
- Re-tested until bulletproof
Eval README (if blind A/B was run):
- States what the data supports AND does not prove
- Records ceiling cells as a positive finding, not noise
- Names the real arbiter (measurement script / production data) if applicable