16 KiB
name, description, triggers, role
| name | description | triggers | role | ||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| browser-debug | Use ONLY when an agent needs to interact with a live browser to debug, verify, or iterate on Web UI work in packages/app — navigate to the dev server, snapshot the DOM/accessibility tree, capture console and network errors, screenshot component states, and exercise user flows end-to-end. Triggers on UI bugs, visual regressions, layout/positioning issues, keyboard-focus problems, console-error reports, and any "open the page and check X" task. Loads the Playwright MCP toolset (browser_navigate, browser_snapshot, browser_take_screenshot, browser_console_messages, browser_evaluate, etc.). Use ONLY for interactive browser work; for writing Playwright *test files* use the existing e2e/ suites under packages/app, not this skill. |
|
Producer |
Core 中立版(Increment 6a 改写,原 deferHard verbatimDir)。机制、结构与 frontmatter 保持;实例术语(工具名、路径、工单号)按
core/adapters/TERMINOLOGY.md绑定到具体实例。
Browser Debug
Interactive browser automation for Web UI work in packages/app. This skill
teaches the recipe for driving the Playwright MCP toolset — when to reach
for it, which tool for which job, and how to capture evidence a reviewer can
trust.
When this skill applies
Reach for browser-debug when a task requires observing the live
application, not just reasoning about code:
- Reproducing a UI bug (focus lost, dropdown won't close, scroll jumps).
- Verifying a visual change actually landed (computed style, layout box, dark-mode token).
- Capturing console errors or failed network requests during a user flow.
- Recording a screenshot / trace as evidence for a review or DoD entry.
- Exercising a keyboard / screen-reader flow end-to-end.
Do not reach for it when:
- You are writing Playwright test files — those live in
packages/app/e2e/...and run viacd packages/app && bun run test:e2e(variant suites:test:e2e:perf,test:e2e:a11y,test:e2e:smoke,test:e2e:visual). This skill is for ad-hoc, agent-driven sessions, not for adding tests. - The question can be answered by reading code or running a unit test. Browser time is expensive; spend it only when the live DOM is the oracle.
Preconditions
- Playwright MCP is configured. The project
octopus.jsonccarries amcp.playwrightblock launching@playwright/mcp. If the MCP tools are not visible, run/kickoffor follow thecodegraph-setupskill's recipe (the closest analog for debugging missing MCP servers). - Dev server is running. The app under test must be reachable. For
packages/app, runbash script/dev-fresh.sh --fullfor the dev server (Windows workstations:bash script/dev-fresh-windows.sh— dev-fresh.sh's stop path depends on setsid/ss, silently ineffective under Git Bash; see local-workstation-quirks.md §5); frontend is athttp://localhost:4444and backend athttp://localhost:4096(see AGENTS.md Dev services section). - Chromium is installed. First run of
@playwright/mcpdownloads it automatically; if the tool reports a missing browser, runbunx playwright install chromiumonce in any shell.
Tool selection recipe
The Playwright MCP exposes many tools. Pick by what you need to observe, not by familiarity:
| Need | Tool | Why |
|---|---|---|
| Open a URL | browser_navigate |
Always the first call. |
| See the page structure (roles, names) | browser_snapshot |
Returns the accessibility tree — stable across CSS refactors. Prefer this over browser_take_screenshot when asserting "does element X exist / is it labelled Y?". |
| See the page visually | browser_take_screenshot |
For layout, spacing, color, overflow. Attach the PNG to your report. |
| Read console errors / warnings | browser_console_messages |
Filter by level (error, warn). The first stop for any "page is blank" bug. |
| Read failed network calls | browser_network_requests |
Locate 4xx/5xx, CORS failures, hanging fetches. |
| Click, type, hover, select | browser_click, browser_type, browser_hover, browser_select_option |
Drive the flow. Always browser_snapshot first to get the ref for the target element. |
| Assert computed style / DOM state | browser_evaluate |
Last resort. Snapshots cover 90%; reach for evaluate only when you need getComputedStyle, scroll positions, or arbitrary JS state. |
| Keyboard navigation | browser_press_key |
For tab order, focus rings, escape-to-close. |
| Record a video / trace | browser_start_tracing, browser_stop_tracing |
Heavy; use only when a bug is timing-sensitive and console+snapshot can't catch it. |
Standard recipe — reproduce a UI bug
Follow these steps in order. Do not skip the snapshot step: the ref it
returns is the stable handle every subsequent tool needs.
browser_navigateto the page where the bug lives.browser_console_messages— record any errors before touching anything. Many "interaction bugs" are actually load-time errors.browser_snapshot— locate the target element; capture itsrefand accessible name.- Drive the interaction with
browser_click/browser_type/browser_press_keyuntil the bug reproduces. - Re-read
browser_console_messagesandbrowser_network_requests— capture the delta introduced by the interaction. browser_take_screenshot(or a secondbrowser_snapshot) to freeze the broken state.- Report: URL, the minimal interaction sequence, console delta, network
delta, screenshot, and the suspected component path in
packages/app/src/....
Standard recipe — verify a UI fix
browser_navigateto the affected page.browser_snapshot— confirm the fixed element is present with the expected accessible name.browser_take_screenshot— full page and element-cropped if the fix is local.- If the fix touched layout,
browser_evaluateagetBoundingClientRect()orgetComputedStyle()on the key node and paste the value into your report. Numbers beat adjectives. - If the fix touched interaction, replay the original bug reproduction steps and assert the previous failure no longer fires.
Standard recipe — capture evidence for verify / DoD
When a verify work item has an E2E or UI DoD entry:
- Run the recipe above for each DoD line.
- Save each screenshot as
.playwright-mcp/dod-{item-id}-{state}.png— the worktree's gitignored output dir. See "Screenshot output paths" below. - In the verification report, cite the file names and tell the reviewer
where to look: open the worktree as an octopus project → session side
panel "Screenshots" tab (or the standalone screenshots page). Verify
with
GET /screenshots?directory={worktree}before citing. A screenshot the reviewer can open beats a paragraph of "I checked it". - When the DoD row will cite the evidence as a
BROWSER:{pack-ref}reference (UI-touching PRs,verifyskill Phase 2.7 / contractbrowser-evidence-4486/shared/evidence-ref-v1), promote the captures into an evidence pack first — feed COPIES throughwriteEvidencePack(<harness-package>/src/browser/evidence-pack.ts, the single mandatory sanitize boundary; the.playwright-mcp/layer keeps its current semantics, additive not replaced) and cite the pack directory. Evidence ladder: raw capture →.playwright-mcp/working layer (gitignored) → sanitized Tier-1 pack →BROWSER:reference the verifier validates mechanically.
Rationalizations to refuse
| Excuse | Reality |
|---|---|
| "The unit test already covers it." | Unit tests assert component logic, not what Chromium actually renders. Visual and focus bugs live in the gap. |
| "I can tell from the JSX that it's aligned." | Computed layout depends on parent containers, fonts, and viewport. Screenshot the real page. |
| "Browser automation is overkill for a small change." | A browser_navigate + browser_take_screenshot pair is ~2 seconds. The cost is in not looking. |
| "I'll just run the whole e2e suite." | The suite answers "did anything break?" — it does not answer "does this specific change look right?". |
| "Snapshot is enough, skip the screenshot." | Accessibility trees hide overflow, clipping, and z-index bugs. Take both for visual changes. |
Boundaries
- Read-only by default. This skill never writes files under
packages/app/src/— it observes. Edits flow throughimplement(bugfix/refactor/port mode). - One page at a time. Close tabs (
browser_close) between scenarios so state doesn't leak. - Evidence lives in the worktree's
.playwright-mcp/dir. Any screenshot or trace you reference in a PR comment, verification report, or DoD entry must be copied into the workflow worktree's.playwright-mcp/directory (gitignored) — never left in/tmp/,playwright-report/, or any machine-local path the reviewer cannot reach through octopus. Cite file names + the worktree project path; the reviewer opens the octopus screenshots UI (session panel tab or screenshots page) to view them.
Session cleanup (mandatory)
A Playwright MCP browser is a real Chrome with a /tmp user-data-dir — its
mode follows mcp.playwright.headless (headless-field-v1, [org-internal #4393]): the repo
default auto spawns headed on desktop (win32/darwin; linux only with
DISPLAY/WAYLAND_DISPLAY set) and headless otherwise; an explicit true/false
or OCTOPUS_MCP_HEADLESS=1|0 pins it. The mode is fixed per new session —
switching the config does not hot-switch a running browser. It
stays alive as long as its MCP server runs, and the MCP server stays alive as
long as its parent octopus backend runs. When a browser-debug session ends
without an explicit browser_close, OR the octopus session times out mid-task,
the Chrome + its MCP server are orphaned and accumulate across sessions —
burning CPU and leaving dead profile dirs (issue [org-internal #1315]); in headed mode the
orphan is a visible stray window.
Iron rule: every browser-debug session MUST close its browser before ending.
- Always close when done. The last action of any browser-debug task is
browser_close. Treat it like areturn— if you opened a browser, you close it. Do not rely on session teardown to do it for you; it doesn't. - Close on error too. If the task aborts (bug not reproduced, tool error,
user redirect), still call
browser_closebefore moving on. An abandoned debugging session is the most common orphan source. - If a browser was already orphaned (you spot a stale Chrome in
ps, or a/tmp/playwright_chromiumdev_profile-*dir with no live session), do not leave it.bash script/dev-fresh.sh --full(Windows workstations:dev-fresh-windows.sh, same stop-path caveat as Preconditions §2) reaps orphan dev sessions and init-adopted browsers before restarting; run it to clean up. For a targeted clean without restarting the dev server, kill the Chrome by its profile:pkill -f 'playwright_chromiumdev_profile'(and anyagent-browserharness reparented to init).
Why this matters: a single forgotten browser_close leaves a Chrome
subtree (8+ processes) running indefinitely. Over days this is hundreds of
orphaned processes and wasted CPU. The MCP server does not auto-close its
browser on disconnect — only an explicit browser_close (or process reaping)
releases it.
Screenshot output paths
The Playwright MCP server is launched WITHOUT --output-dir (see
octopus.jsonc → mcp.playwright.command): its default output location is
.playwright-mcp/ in the process cwd (the worktree root), gitignored and
auto-evicted at 100 MB via --output-max-size. Auto-named output (no
filename) lands there — but an explicit filename can resolve against the
cwd and drop the PNG into the repo root, where the /*.png gitignore
backstop hides it from git status and nothing ever cleans it up.
Iron rule: 截图文件名必须写进 .playwright-mcp/。 Every filename
passed to browser_take_screenshot (or any tool that writes a file) MUST
resolve into .playwright-mcp/ — write it as .playwright-mcp/<name>.png,
or use an absolute path under /tmp/octopus/. NEVER a bare relative name
like 2202-after-models.png: it leaks a root-level PNG that gitignore masks
but nobody deletes (the 2026-08 repo-root buildup — 40+ orphan PNGs — came
from exactly this).
Two valid destinations for screenshots, by purpose:
| Purpose | Destination | Why |
|---|---|---|
| Ad-hoc / throwaway (debug a layout, confirm a fix landed, never cited in any report) | .playwright-mcp/<name>.png, or omit filename to auto-save into the same dir |
Survives across the session, auto-evicted at 100 MB, never pollutes the repo. |
| Evidence cited in a PR / report / DoD entry | .playwright-mcp/<name>.png — the cwd is the worktree root, so no copy step is needed |
Zero repo bloat (gitignored, never committed); viewable by the reviewer through octopus's own screenshots UI (session panel "Screenshots" tab / standalone screenshots page, served by GET /screenshots?directory={worktree}); lifecycle bounded to the worktree = the review window. |
Evidence publishing ladder (verified 2026-08-14 on this Gitea version):
- Copy into the workflow worktree's
.playwright-mcp/— primary path. Cite file names + worktree project path; reviewer views via the octopus screenshots UI. Verify reachability withGET /screenshots?directory={worktree}before citing. Never commit evidence binaries to git — screenshots accumulate per PR and git history is unreclaimable (measured ~80KB/PR today, unbounded growth). Do not put routine debug screenshots on the wiki either — base64-in-page bloats the wiki git repo equivalently; the wiki evidence namespace is reserved for the durable verify-stage evidence channel (rung 3). - Issue attachments API (
POST /api/v1/repos/{o}/{r}/issues/{n}/assets) — alternative when the reviewer cannot reach the worktree through the octopus UI. The endpoint exists but has returned 500 on this instance; verify before relying on it. - Publish to the wiki evidence namespace
{slug}/verify/evidence/{name}viawiki 读写 API(见 TERMINOLOGY)— last resort only, when no other channel works AND the evidence must outlive the worktree (this is the same durable channelverifymandates for DoD evidence entries — seecore/skills/verify/SKILL.mdReferences). Keep it small: ≤5 images per PR, each ≤200KB (compress/downscale, prefer element crops over full-page).
Iron rule: never cite a machine-local path (/tmp/..., ~, project
root) as evidence in a Tier-2 output (PR comment, issue comment,
verification report). The reviewer must be able to open the evidence
through octopus (.playwright-mcp/ + screenshots UI) or a repo/wiki URL
(worktree files, or the {slug}/verify/evidence/{name} wiki pages). A
citation they cannot reach is a fabricated citation.
Never write screenshots to the project root. The .gitignore rejects
/*.png, /.playwright-mcp/, and /playwright-report/ only as a backstop —
an ignored file is still an orphan on disk. The .playwright-mcp/ filename
prefix is the real fix; rely on that, not on the backstop.
References
core/skills/frontend/SKILL.md— where UI changes are planned; invoke this skill during Phase 3 / Phase 4 to verify each state visually.core/skills/implement/SKILL.md— Frontend Mode; pair withbrowser-debugwhen the work item touches rendering.core/skills/implement/SKILL.md(Mode: bugfix) — Phase 1 reproduction; use this skill to capture the failing state before isolating root cause.core/skills/verify/SKILL.md— Phase 2 / DoD matrix; use this skill to generate screenshot evidence for UI-tagged DoD entries.packages/app/e2e/— permanent Playwright test suites; this skill is the ad-hoc complement, not a replacement.rules/dev-serveron the wiki — how to run the dev server for TUI work; consultpackages/app/README.mdfor the web equivalent.