Files
octopus-workflow/core/skills/image-interpret/SKILL.md
T

2.9 KiB

name, description, triggers, role
name description triggers role
image-interpret Use ONLY when the agent needs to read or analyze an image (screenshot, design mockup, UI capture, .png/.jpg/.webp file) but the active model lacks vision capability. Delegates to the Image Interpreter role (vision-capable model) which returns a structured text spec — layout regions, colors (hex), spacing (px), typography, component hierarchy, and responsive breakpoints. Do not use for non-image tasks or when the active model already supports vision.
screenshot
mockup
design mockup
UI capture
.png
.jpg
.jpeg
.webp
image
读图
设计稿
截图
视觉
Tool

Image Interpret — delegate image reading to a vision-capable role

When to use

You (the calling agent) cannot read images — your model has no vision. Any task that requires seeing a visual artifact must be delegated to the Image Interpreter role, which runs on a vision-capable model and returns a structured text description you can act on.

Trigger situations:

  • The user references a screenshot, design mockup, or image file.
  • A task requires comparing rendered UI against a design.
  • You need layout/color/spacing specs extracted from a picture.
  • A file path ending in .png / .jpg / .jpeg / .webp must be understood.

Workflow

  1. Identify the image. Confirm the file path (or paths) that must be read. Resolve it to an absolute path the subagent can access.

  2. Frame the question. Decide exactly what the caller needs — full layout spec, a specific region, color tokens, component list, diff vs. existing UI.

  3. Delegate via the Task tool with role: "Tool". Do NOT pass subagent_type — it is derived automatically from the role.

    Task(
      role: "Tool",
      description: "Extract header layout spec",
      prompt: "Read the image at /abs/path/mockup.png. Extract: (1) header
               layout regions with bounding positions, (2) all colors as hex,
               (3) spacing in px, (4) typography (family/size/weight). Return
               a structured text spec the caller can implement without seeing
               the image."
    )
    
  4. Consume the returned text spec. The Image Interpreter returns a structured description only — it cannot write code or artifacts. Use the returned spec to drive your own work (code generation, design review, planning).

  5. Do not attempt to read the image yourself. If you are a non-vision model, opening the image yields nothing. Always delegate.

Notes

  • The Image Interpreter is read-only — it describes images, it does not edit files or write code.
  • One delegation per distinct visual question. Batch multiple images into a single delegation only when they must be compared.
  • If the returned spec is ambiguous, delegate again with a sharper question rather than guessing.