2.9 KiB
name, description, triggers, role
| name | description | triggers | role | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| image-interpret | Use ONLY when the agent needs to read or analyze an image (screenshot, design mockup, UI capture, .png/.jpg/.webp file) but the active model lacks vision capability. Delegates to the Image Interpreter role (vision-capable model) which returns a structured text spec — layout regions, colors (hex), spacing (px), typography, component hierarchy, and responsive breakpoints. Do not use for non-image tasks or when the active model already supports vision. |
|
Tool |
Image Interpret — delegate image reading to a vision-capable role
When to use
You (the calling agent) cannot read images — your model has no vision. Any task that requires seeing a visual artifact must be delegated to the Image Interpreter role, which runs on a vision-capable model and returns a structured text description you can act on.
Trigger situations:
- The user references a screenshot, design mockup, or image file.
- A task requires comparing rendered UI against a design.
- You need layout/color/spacing specs extracted from a picture.
- A file path ending in
.png/.jpg/.jpeg/.webpmust be understood.
Workflow
-
Identify the image. Confirm the file path (or paths) that must be read. Resolve it to an absolute path the subagent can access.
-
Frame the question. Decide exactly what the caller needs — full layout spec, a specific region, color tokens, component list, diff vs. existing UI.
-
Delegate via the Task tool with
role: "Tool". Do NOT passsubagent_type— it is derived automatically from the role.Task( role: "Tool", description: "Extract header layout spec", prompt: "Read the image at /abs/path/mockup.png. Extract: (1) header layout regions with bounding positions, (2) all colors as hex, (3) spacing in px, (4) typography (family/size/weight). Return a structured text spec the caller can implement without seeing the image." ) -
Consume the returned text spec. The Image Interpreter returns a structured description only — it cannot write code or artifacts. Use the returned spec to drive your own work (code generation, design review, planning).
-
Do not attempt to read the image yourself. If you are a non-vision model, opening the image yields nothing. Always delegate.
Notes
- The Image Interpreter is read-only — it describes images, it does not edit files or write code.
- One delegation per distinct visual question. Batch multiple images into a single delegation only when they must be compared.
- If the returned spec is ambiguous, delegate again with a sharper question rather than guessing.