Initial publish v0.1.0: standalone workflow core (corpus + examples + guards)
This commit is contained in:
@@ -0,0 +1,82 @@
|
||||
---
|
||||
name: image-interpret
|
||||
description: >
|
||||
Use ONLY when the agent needs to read or analyze an image (screenshot,
|
||||
design mockup, UI capture, .png/.jpg/.webp file) but the active model lacks
|
||||
vision capability. Delegates to the Image Interpreter role (vision-capable
|
||||
model) which returns a structured text spec — layout regions, colors (hex),
|
||||
spacing (px), typography, component hierarchy, and responsive breakpoints.
|
||||
Do not use for non-image tasks or when the active model already supports
|
||||
vision.
|
||||
triggers:
|
||||
- screenshot
|
||||
- mockup
|
||||
- design mockup
|
||||
- UI capture
|
||||
- .png
|
||||
- .jpg
|
||||
- .jpeg
|
||||
- .webp
|
||||
- image
|
||||
- 读图
|
||||
- 设计稿
|
||||
- 截图
|
||||
- 视觉
|
||||
role: Tool
|
||||
---
|
||||
|
||||
# Image Interpret — delegate image reading to a vision-capable role
|
||||
|
||||
## When to use
|
||||
|
||||
You (the calling agent) cannot read images — your model has no vision. Any
|
||||
task that requires *seeing* a visual artifact must be delegated to the
|
||||
**Image Interpreter** role, which runs on a vision-capable model and returns a
|
||||
structured text description you can act on.
|
||||
|
||||
Trigger situations:
|
||||
|
||||
- The user references a screenshot, design mockup, or image file.
|
||||
- A task requires comparing rendered UI against a design.
|
||||
- You need layout/color/spacing specs extracted from a picture.
|
||||
- A file path ending in `.png` / `.jpg` / `.jpeg` / `.webp` must be understood.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. **Identify the image.** Confirm the file path (or paths) that must be read.
|
||||
Resolve it to an absolute path the subagent can access.
|
||||
|
||||
2. **Frame the question.** Decide exactly what the caller needs — full layout
|
||||
spec, a specific region, color tokens, component list, diff vs. existing UI.
|
||||
|
||||
3. **Delegate via the Task tool with `role: "Tool"`.** Do NOT pass
|
||||
`subagent_type` — it is derived automatically from the role.
|
||||
|
||||
```
|
||||
Task(
|
||||
role: "Tool",
|
||||
description: "Extract header layout spec",
|
||||
prompt: "Read the image at /abs/path/mockup.png. Extract: (1) header
|
||||
layout regions with bounding positions, (2) all colors as hex,
|
||||
(3) spacing in px, (4) typography (family/size/weight). Return
|
||||
a structured text spec the caller can implement without seeing
|
||||
the image."
|
||||
)
|
||||
```
|
||||
|
||||
4. **Consume the returned text spec.** The Image Interpreter returns a
|
||||
structured description only — it cannot write code or artifacts. Use the
|
||||
returned spec to drive your own work (code generation, design review,
|
||||
planning).
|
||||
|
||||
5. **Do not attempt to read the image yourself.** If you are a non-vision
|
||||
model, opening the image yields nothing. Always delegate.
|
||||
|
||||
## Notes
|
||||
|
||||
- The Image Interpreter is **read-only** — it describes images, it does not edit
|
||||
files or write code.
|
||||
- One delegation per distinct visual question. Batch multiple images into a
|
||||
single delegation only when they must be compared.
|
||||
- If the returned spec is ambiguous, delegate again with a sharper question
|
||||
rather than guessing.
|
||||
Reference in New Issue
Block a user