Initial publish v0.1.0: standalone workflow core (corpus + examples + guards)
This commit is contained in:
@@ -0,0 +1,260 @@
|
||||
---
|
||||
name: writing-skills
|
||||
description: Use ONLY when creating, editing, or evaluating an octopus skill — writing the SKILL.md body, deciding whether a skill needs baseline testing, or running a blind A/B eval of a behavior-shaping technique. Use BEFORE first deployment of any skill that must hold a process gate under pressure.
|
||||
triggers:
|
||||
# High-discrimination phrases only ([org-internal #4525]). We deliberately do NOT use:
|
||||
# - bare "skill" / 技能 — collides with every skill discussion. Worse: the
|
||||
# hard skill gate folds the previous assistant turn into its match text on
|
||||
# approval turns, so any model reply containing 技能 (e.g. a gate
|
||||
# disclaimer) re-fired the gate every turn — a self-sustaining loop.
|
||||
# - this skill's own name — the gate reminder lists matched skill names, so
|
||||
# a model echoing them re-seeds the trigger for the next turn.
|
||||
# - "skill gate" — names the mechanism, not the task of authoring skills.
|
||||
# Note matches() is a CONTIGUOUS substring, so verb phrases like
|
||||
# "evaluate skill" miss real input "evaluate THIS skill". CJK recall is
|
||||
# accepted as imperfect (写个技能 breaks 写技能 contiguity): the skills
|
||||
# listing's description remains the model-side routing path for prose asks.
|
||||
- blind A/B
|
||||
- A/B eval
|
||||
- skill eval
|
||||
- 技能编写
|
||||
- 编写技能
|
||||
- 技能评估
|
||||
- 评估技能
|
||||
- 行为塑造
|
||||
- SKILL.md
|
||||
role: Producer
|
||||
---
|
||||
|
||||
# Writing Skills
|
||||
|
||||
## Overview
|
||||
|
||||
**Writing a skill IS test-driven development applied to process documentation.**
|
||||
You write test cases (blind pressure scenarios run through Explorer subagents),
|
||||
watch them fail (baseline behavior without the skill), write the skill, watch
|
||||
tests pass (agents comply), and refactor (close loopholes).
|
||||
|
||||
**Core principle:** If you didn't watch an agent fail without the skill, you
|
||||
don't know if the skill teaches the right thing — or teaches anything at all.
|
||||
|
||||
This skill covers the *methodology*. For file locations, frontmatter fields,
|
||||
and valid config shapes, load `customize-octopus` — a built-in skill
|
||||
(`packages/octopus/src/skill/prompt/customize-octopus.md`), not a skill under
|
||||
`.octopus/skills/`. The two co-trigger: mechanics + methodology.
|
||||
|
||||
## The Iron Law
|
||||
|
||||
```
|
||||
NO BEHAVIOR-SHAPING SKILL WITHOUT A FAILING BASELINE FIRST.
|
||||
```
|
||||
|
||||
This applies to NEW skills AND EDITS that change what a skill enforces. Write
|
||||
the skill before testing the baseline? You no longer have a baseline — the
|
||||
failure you needed to document is already contaminated by knowing the fix.
|
||||
|
||||
No exceptions for "simple additions", "just a new section", or "documentation
|
||||
updates". If the edit changes agent behavior, it needs a baseline.
|
||||
|
||||
## When to eval — and when it is overkill (decide FIRST)
|
||||
|
||||
Not every skill needs a blind A/B eval. Classify before authoring:
|
||||
|
||||
| Skill type | Example | Eval required? |
|
||||
|---|---|---|
|
||||
| **Behavior-shaping / discipline** (must hold a gate under pressure) | `review-code` (Phase C), a new "never skip X" rule | **YES — full blind A/B** (RED-GREEN-REFACTOR below) |
|
||||
| **Technique / how-to** (a method agents apply) | `condition-based-waiting`, `root-cause-tracing` port | Application scenarios (can the agent apply it correctly?) — baseline optional |
|
||||
| **Reference / mechanical** (lookup, syntax, config shapes) | config-field docs, command references | **NO** — source-fidelity check + lint suffice |
|
||||
| **Pure port of a proven technique** (verbatim, minimal localization) | a 1:1 port with wording unchanged | Fidelity check against source + spot-check; full re-eval only if localized substantially |
|
||||
|
||||
**The question that decides it:** "Could a capable agent comply with this skill
|
||||
*without ever reading it*?" If yes (it restates convention, or the gate is
|
||||
already procedurally explicit elsewhere), the skill is reference-class and eval
|
||||
will show a **ceiling** — both control and treatment hold, and the eval carries
|
||||
no signal. Do not run a blind eval expecting to prove a reference skill works;
|
||||
run a fidelity check instead.
|
||||
|
||||
## RED-GREEN-REFACTOR for skills
|
||||
|
||||
### RED — write the failing test (baseline)
|
||||
|
||||
Run pressure scenarios through an Explorer subagent **WITHOUT** the skill (or
|
||||
with a control SKILL that omits only the technique under test). Document, verbatim:
|
||||
- What choices the agent made.
|
||||
- The exact rationalizations it used to skip the gate.
|
||||
- Which pressures (deadline, sunk cost, authority, exhaustion) triggered the violation.
|
||||
|
||||
This is "watch the test fail". You must see what agents naturally do before
|
||||
writing the skill, or you are writing to an imagined failure.
|
||||
|
||||
### GREEN — write the minimal skill
|
||||
|
||||
Write the skill to address **the specific rationalizations you observed** — not
|
||||
hypothetical ones. Run the same scenarios WITH the skill. The agent should now
|
||||
comply. Do not add content for failures you did not see.
|
||||
|
||||
### REFACTOR — close loopholes
|
||||
|
||||
The agent found a new rationalization in GREEN? Add an explicit counter (see
|
||||
Rationalization Table below). Re-test until bulletproof.
|
||||
|
||||
## Blind A/B eval methodology
|
||||
|
||||
For behavior-shaping skills, the eval is a **blind controlled experiment**:
|
||||
control = SKILL without the technique; treatment = SKILL with it. The
|
||||
Explorer subagents running the scenarios are unaware of the condition **or**
|
||||
that they are being tested.
|
||||
|
||||
### Artifact layout (wiki pages under `_evals/{skill-name}/`)
|
||||
|
||||
> The `_evals/{skill-name}/` namespace is an intentional slug-less collection
|
||||
> namespace — eval records aggregate across skills and outlive any one
|
||||
> pipeline run — analogous to two documented slug-convention exceptions: the
|
||||
> audit `audit/{date}/` date-slug exception (NAM 4.4,
|
||||
> `core/checklists/audit-process.md`) and the retrospective
|
||||
> `_retrospectives/` exception (`retrospective/SKILL.md` Phase 7 — Report).
|
||||
|
||||
```
|
||||
README findings, conclusion, honest evidence strength,
|
||||
and why you stopped (or didn't)
|
||||
round-1-initial/control-SKILL baseline skill (technique omitted)
|
||||
round-1-initial/treatment-SKILL skill with the technique
|
||||
round-1-initial/scenarios/scenario-<id> blind pressure scenarios (3+ for discipline skills)
|
||||
round-1-initial/cells/cell-<id>-{control,treatment} per-cell dispatch prompts
|
||||
round-1-initial/outputs/output-<id>-{control,treatment} captured agent outputs (verbatim)
|
||||
round-1-initial/rubric scoring rubric
|
||||
round-1-initial/scoring scored results table
|
||||
round-2-replication/* re-runnable half; same names, NEW scenarios
|
||||
```
|
||||
|
||||
### Procedure
|
||||
|
||||
1. **Write 3+ pressure scenarios** that combine pressures (deadline + sunk cost;
|
||||
authority + exhaustion). Each scenario tempts the agent to skip the gate the
|
||||
skill enforces.
|
||||
2. **Run each scenario twice, blind**, via an Explorer subagent: once seeded
|
||||
with `control-SKILL.md`, once with `treatment-SKILL.md`. The subagent sees
|
||||
only its assigned SKILL + the scenario — never the other condition, never
|
||||
the word "test" or "experiment".
|
||||
3. **Score outputs against the rubric** on a 3-level verdict: `HOLD` (kept the
|
||||
gate), `PARTIAL` (hedged / delayed but technically held), `RATIONALIZE`
|
||||
(skipped with justification).
|
||||
4. **Treatment must win the *differentiating* cells** — cells where control and
|
||||
treatment split. **Ceiling cells** (both HOLD) carry NO signal either way:
|
||||
they are a positive finding (that gate is robust without the technique), not
|
||||
noise to eliminate by cranking pressure.
|
||||
5. **Replicate**: run a second round with NEW scenarios before concluding. A
|
||||
single round is one data point.
|
||||
6. **Record honest evidence strength** in the README — see calibration below.
|
||||
|
||||
### Worked example
|
||||
|
||||
The canonical blind A/B run (2 rounds, 17 cells, rationalization-anticipation
|
||||
tables) is published at the `_evals/skill-compliance-framing/` wiki namespace
|
||||
once its artifacts are promoted from internal storage. The experiment
|
||||
demonstrated a ceiling finding: the tested gate was robust in both control
|
||||
and treatment arms without the technique under evaluation.
|
||||
|
||||
## Honest evidence calibration (write this in the eval README)
|
||||
|
||||
An eval README that claims more than the data supports is worse than no eval.
|
||||
State explicitly:
|
||||
|
||||
- **Supports**: the construct the data is consistent with (e.g. "*consistent
|
||||
with* a real but **construct-specific** effect on redundancy-gates under
|
||||
pressure").
|
||||
- **Does NOT prove**: the limits — small differentiating N, ceiling cells,
|
||||
stated-decision eval vs production compliance.
|
||||
- **Ceiling IS a finding**: when both groups hold a gate, that gate is robust
|
||||
without the technique. Do not manufacture pressure to break the control —
|
||||
that is eval p-hacking.
|
||||
- **Real arbiter**: for production skills, the honest next test is often a
|
||||
measurement script over real iterations (e.g. `script/measure-review-blockers.ts`),
|
||||
not more simulated cells.
|
||||
|
||||
**Stopping is a valid decision.** When the ceiling pattern is clear and
|
||||
differentiating N stays small, further stated-decision cells have diminishing
|
||||
returns. Say so and stop.
|
||||
|
||||
## Match the form to the failure
|
||||
|
||||
Before authoring guidance, classify the baseline failure. The form that
|
||||
bulletproofs one failure type measurably backfires on another.
|
||||
|
||||
| Baseline failure | Right form | Wrong form |
|
||||
|---|---|---|
|
||||
| Skips/violates a rule under pressure (knows better, does it anyway) | Prohibition + rationalization table + red flags | Soft guidance ("prefer...", "consider...") |
|
||||
| Complies, but output has the wrong shape (bloated prompt, buried verdict) | Positive recipe: state what the output IS — its parts, in order | Prohibition list ("don't restate", "never narrate") |
|
||||
| Omits a required element from something they already produce | Structural: REQUIRED field or slot in the template | Prose reminders near the template |
|
||||
| Behavior should depend on a condition | Conditional keyed to an observable predicate ("if the brief exists, reference it") | Unconditional rule + exemption clauses |
|
||||
|
||||
**Why prohibitions backfire on shaping problems:** under a competing incentive,
|
||||
agents negotiate with "don't X". In head-to-head wording tests, the prohibition
|
||||
arm produced more of the unwanted content than the recipe arm — and trended
|
||||
worse than even the no-guidance control. A recipe leaves nothing to negotiate:
|
||||
the output matches the stated shape or it doesn't.
|
||||
|
||||
**Rules for whichever form:**
|
||||
- **No nuance clauses.** "Don't X unless it matters" reopens the negotiation.
|
||||
Express a real exception as its own conditional on an observable predicate.
|
||||
- **Exemption clauses don't scope.** "This limit doesn't apply to code blocks"
|
||||
still suppresses code blocks. Restructure so the rule can't reach the exempt part.
|
||||
|
||||
## Rationalization table
|
||||
|
||||
Capture the rationalizations you observed in RED. Every excuse the baseline
|
||||
agents used goes in the table, paired with the reality that defeats it:
|
||||
|
||||
```markdown
|
||||
| Excuse | Reality |
|
||||
|--------|---------|
|
||||
| "Too simple to test" | Simple code breaks. The test takes 30 seconds. |
|
||||
| "I'll test after" | Tests-after answers "what does this do?"; the baseline answers "what SHOULD it do?" |
|
||||
```
|
||||
|
||||
This is the same rationalization-anticipation device already embedded in
|
||||
`implement` (including its bugfix mode) — here it is an *authoring* tool: you build it from
|
||||
your own baseline data, not from guesswork.
|
||||
|
||||
## Micro-test wording before full scenarios
|
||||
|
||||
Full blind A/B runs are expensive. Verify the wording itself first with
|
||||
micro-tests:
|
||||
|
||||
1. **One fresh-context sample per call** — a single-shot Explorer subagent.
|
||||
System prompt = the realistic context the guidance will live in; user
|
||||
message = a task that tempts the failure.
|
||||
2. **Always include a no-guidance control.** If the control doesn't exhibit the
|
||||
failure, there is nothing to fix — stop, don't author the guidance.
|
||||
3. **5+ reps per variant.** Single samples lie.
|
||||
4. **Read every flagged match manually.** Template echoes and quoted
|
||||
counter-examples masquerade as hits; automated counts overstate both failure
|
||||
and success.
|
||||
5. **Variance is a metric.** When wording lands, reps converge on the same
|
||||
shape. Five different interpretations across five reps means the wording
|
||||
isn't binding — tighten the form before adding words.
|
||||
|
||||
Micro-tests verify wording; they do not replace blind A/B for discipline skills.
|
||||
|
||||
## Authoring checklist
|
||||
|
||||
**RED — baseline:**
|
||||
- [ ] Classified the skill (behavior-shaping / technique / reference / port)
|
||||
- [ ] If behavior-shaping: wrote 3+ combined-pressure scenarios
|
||||
- [ ] Ran baseline (control SKILL) blind via Explorer subagent; documented rationalizations verbatim
|
||||
|
||||
**GREEN — minimal skill:**
|
||||
- [ ] Frontmatter `name` + `description` (description = WHEN to use, not what it does — see `customize-octopus`)
|
||||
- [ ] Addresses the specific baseline failures observed (not hypothetical ones)
|
||||
- [ ] Form matches the failure type (table above)
|
||||
- [ ] Ran scenarios WITH skill; agent now complies
|
||||
|
||||
**REFACTOR — loopholes:**
|
||||
- [ ] New rationalizations from testing have explicit counters
|
||||
- [ ] Rationalization table built from real baseline data
|
||||
- [ ] Re-tested until bulletproof
|
||||
|
||||
**Eval README (if blind A/B was run):**
|
||||
- [ ] States what the data supports AND does not prove
|
||||
- [ ] Records ceiling cells as a positive finding, not noise
|
||||
- [ ] Names the real arbiter (measurement script / production data) if applicable
|
||||
Reference in New Issue
Block a user