Skip to content

Design Your AI Workflow

Part of: AI Workflow Framework

The Design phase is where you decide how your workflow should be built — before you build it. You take the Workflow Requirements from the Deconstruct step (the what) and produce a Design Spec (the how) — the architectural blueprint Build consumes to generate skills, agents, and prompts.

The Workflow Requirements is the canonical source of truth. The Design Spec references it, never restates it. Build reads both together.

Design walks through three layers of decisions that build on each other. Each layer is approved before the next begins, so strategic errors get caught before you’ve invested time in detailed specs:

LayerWhat it decidesWhy it’s separate
Layer 1 — ArchitecturePlatform, mechanism (Skill/Agent), autonomy level, packaging, model class, integration optionsStrategic. Cheap to revisit. Wrong call here cascades everywhere.
Layer 2 — DecompositionFor each step (or capability domain), what AI building block delivers it: a new skill, an existing skill (as-is or extended), an inline prompt block, an agent, or a human actionStructural. Decides what gets built and what gets reused.
Layer 3 — Component BlueprintsField-level specs for each new skill and agent — name, description, inputs/outputs, decision logic, failure modes, tools, deploymentDetailed. Most expensive to redo. Build uses these to generate artifacts.

The Design skill walks you through these layers in order, with a lightweight confirmation moment between each. Approval of the draft spec is the only hard gate.

What you’ll doConfirm your platform, work through three layers of design decisions with lightweight confirmation at each handoff, and approve the final spec
What you’ll getA Design Spec — three-layer architecture and component blueprints, with frontmatter, stable IDs, and a Self-Test Summary
Time~30 minutes

Not every workflow needs the same level of AI infrastructure. A weekly status report is a skill you start by name that follows the same steps every time. A multi-department content pipeline may need an agent that decides its own path as it goes. Choosing the wrong mechanism means either over-engineering (an agent where a skill would do) or under-building (a skill forced to make agent-level decisions).

Design also makes the workflow portable. Because the Design Spec follows the agentskills.io standard and uses stable IDs for every component, the same spec can be built on Claude Code, Claude.ai, Cowork, Codex, ChatGPT, or Gemini CLI — with only the platform-specific bits (Packaging, Deployment Plan) adjusted at the handoff.


Strategic decisions that shape everything downstream. You start here.

You name your AI platform — at whatever level of specificity you have. “Claude Code,” “ChatGPT,” “Google Gemini,” “Claude” are all fine. The ecosystem is enough for Design decisions; the specific offering (Claude Code vs. Claude.ai vs. Cowork) is resolved when Build generates artifacts.

Other Architecture Decisions extracted from the Workflow Requirements

Section titled “Other Architecture Decisions extracted from the Workflow Requirements”

Rather than walking through a checklist, the Design skill uses an extract-then-confirm approach: ask one question (platform), extract everything else from the Workflow Requirements, present the analysis for confirmation.

  • Tool integrations — pulled from per-step Inputs, Context Needed, and the Context Inventory
  • Trigger/schedule — pulled from the Metadata table; time-based triggers imply scheduled execution
  • Context readiness flags — pulled from the Context Inventory’s AI Accessible column; flags items needing resolution before Build

Where the whole workflow sits on the autonomy spectrum. Autonomy asks how much the AI decides on its own — look at what decides the next step:

Human ———— Deterministic ———————— Guided ———————— Autonomous
(a person does it) (you give instructions) (bounded decisions, your method) (you give a goal)
LevelSignalsOrchestration implications
HumanA person performs the step — judgment, creativity, approval, or physical action; no AINo AI artifact — captured as Human step in the Decomposition table
DeterministicYou set every step and the AI carries each one out. It may write or summarize inside a step, but its output never changes what happens next. Test: does the work follow the same path whatever the AI produces?Skill likely sufficient
GuidedYou set the structure and the methodology (a rubric, criteria, a process); the AI uses it to make bounded decisions on your behalf — route an item, choose a tool, judge quality and send work back. Test: does the AI’s judgment, made by your rules, decide what happens next?Skill or Agent
AutonomousThe AI plans its own steps, decides what to do next at each turn, and keeps going until the goal is met; its decision-making is open-ended. Test: could you only describe the goal, not the steps?Agent required

Writing or summarizing inside a step never makes a workflow Guided, and the number of agents doesn’t set the level. A branch on a value the AI didn’t judge — an API’s score, a timer, a field value — is still an instruction: Deterministic. A person approving the AI’s decisions doesn’t lower the level. The workflow’s level is the highest level any of its AI steps reaches. Whether a person should check the AI’s output is answered by the involvement mode below and by Test, not by autonomy. The Guided/Autonomous line follows Anthropic’s Building Effective Agents: workflows run on “predefined code paths” (routing and evaluator-optimizer loops are Guided here), while agents “dynamically direct their own processes and tool usage” (Autonomous).

Based on the autonomy assessment and architecture decisions, the model recommends who drives the workflow:

MechanismDescriptionSignals
SkillYou start it by name; it follows the mapped steps, pausing where you saidYou trigger the work; same steps each time; decisions are yours at the pauses
AgentDecides its own path at runtime, uses tools on its judgment, can run unattendedSteps depend on what it finds; scheduled or hands-off runs

Plus an involvement mode:

ModeDescriptionDetermined by
AugmentedA person is in the workflow along the way, guiding, engaging, or collaborating with the AI while it runsA review or approval pause, questions the AI asks mid-run, or a person working alongside it in the session
AutomatedNo one takes part until it’s doneNobody takes part between start and finish — scheduled, triggered, or started by hand (starting a run doesn’t make it Augmented)

How the artifacts ship together:

ValueWhen to use
PluginMultiple related artifacts shipped together as a marketplace plugin (e.g., handsonai-plugins layout)
Standalone SkillA single skill, uploaded directly (zip for Claude.ai, single SKILL.md for code-mode platforms, single skill for ChatGPT)
Workspace AgentA ChatGPT Workspace Agent bundling orchestration + skills + tools (the current ChatGPT primitive; Custom GPTs are deprecated)
Loose FilesIndividual files in a project directory, no distribution wrapper

A capability tier (reasoning-heavy / balanced / fast / vision) with per-step overrides if needed. The spec records only the tier — Build resolves the concrete model name for the target platform via web search at generation time, so specs never carry model IDs that go stale.

For each tool the workflow needs, the model checks for a platform-native connector first, then falls back to MCP servers, APIs, SDKs, and CLIs — with source URLs and trade-offs. The output is platform-agnostic; Build does the per-platform setup research.

Before Layer 1 closes, the design answers four questions in plain language — they matter most for workflows that write to live systems, run unattended, or process content you didn’t author:

  1. Write access — Which connected tools can this workflow create, modify, or send through? Apply least privilege: only the scopes it needs, and prefer draft-don’t-send until trust is established.
  2. Untrusted input — Does any step process content you didn’t author (inbound email, web pages, form submissions)? That content must be treated as data, never as instructions — this is how prompt-injection incidents happen.
  3. Unattended runs — Scheduled or headless? Then human gates on outward-facing actions, a cap on actions per run, and a log of every write.
  4. Blast radius — What’s the worst realistic outcome of a bad run? Put a human gate in front of that action.

The answers and mitigations are recorded in the spec’s Safety & Permissions section; Build enforces them during connector setup, and Run re-verifies them before the first scheduled run. For a read-only, human-triggered workflow this is one sentence, not a hurdle.

End of Layer 1. Lightweight confirmation: “Architecture confirmed: [summary]. Moving to Decomposition. Confirm to proceed.”


For each step (or capability domain), decide which AI building block delivers it.

Steps are defined in the Workflow Requirements. This table adds the building-block classification and the concrete Build output for each:

ColumnValues
StepStep ID from Workflow Requirements
AutonomyHuman / Deterministic / Guided / Autonomous
OrchestrationPrompt / Skill / Agent
IntegrationBlock + tool + use/build tag (e.g., “MCP: HubSpot (use)“)
IntelligenceModel class + context source IDs + memory flag
Build OutputOne of: New skill: S1 / Use existing: [name] / Extend existing: [name] / New agent: A1 / Inline prompt → Workflow Requirements Step N / Handled by orchestrator / MCP server: [name] / Human (no artifact)
Human Gate?Yes / No (from Workflow Requirements Human Gates)

For goal-driven workflows, this is replaced by a Capability Domain Mapping — capability domains are derived during Design (not present in the Workflow Requirements). A capability domain is a durable competency the agent draws on (e.g., “research,” “synthesis”) — not a step or pipeline stage. Collapse parallel applications of one competency into a single domain (expressed as a fan-out rather than duplicate rows), and treat domains as capabilities available to the orchestrator at runtime, not a fixed path. Each maps to integration needs, intelligence requirements, and a Build Output.

When mechanism is Skill, the spec includes an Orchestrator Prompt Outline — the structural skeleton of the workflow’s main prompt. It names which step invokes which skill, where PAUSE points sit (from Human Gates), what the user provides at each gate, and — as its last element — the closing What I did run summary the orchestrator gives the user at the end of every run. Build expands the outline into the full orchestrator using Step Details from the Workflow Requirements.

Omitted for Agent — orchestration logic lives in the Deployment Plan; on Claude Code and Cowork the primary session orchestrates and the agents are its workers.

Context items from the Workflow Requirements’ Context Inventory flagged as Partial or No for AI Accessible — with the action needed to make each one accessible before Build runs.

Quick Wins → Core → Future Enhancement. Within each tier, dependencies follow the Depends On field of each skill.

End of Layer 2. Lightweight confirmation: “Decomposition confirmed: [counts]. Moving to Component Blueprints. Confirm to proceed.”


Field-level specs Build uses to generate each new skill and agent.

S1 is always the orchestrator skill for a Skill mechanism — it carries the workflow’s name; component skills follow.

For each step tagged New skill: SN:

FieldPurpose
ID, Name, DescriptionIdentity. Name must be lowercase-hyphenated, ≤64 chars, and capability-named (gerund/verb-object like summarizing-transcripts, never workflow-coupled — only the orchestrator skill takes the workflow name). Description must start with “This skill should be used when…”, be ≤1024 chars, be written in third person, and name concrete trigger keywords — this is the verbatim text that goes into the SKILL.md frontmatter and drives auto-activation.
Purpose, Covers Steps/DomainsInternal context for spec readers
Inputs, OutputsThe skill’s contract
Decision Logic, Failure ModesWhat the skill does and how it handles edge cases
Required Tools, Depends OnRuntime dependencies on tools and other skills
Stateful?Memory building-block flag

Agent Configuration (14 fields, mandatory when mechanism = Agent)

Section titled “Agent Configuration (14 fields, mandatory when mechanism = Agent)”

For each step tagged New agent: AN:

FieldPurpose
ID, Name, DescriptionIdentity. Description must start with “Use this agent when…” and is the verbatim text for the agent file frontmatter.
Mission, Responsibilities, Output Format, Tone & Style, ConstraintsThe agent’s behavior — decomposed into structured sub-fields, not jammed into a single “Instructions” cell. For orchestrator-dispatched workers, Output Format is the handoff contract; for Autonomous agents, Constraints includes an iterations-per-run bound.
Failure ModesCondition → action, one per line — including what the agent returns to its orchestrator when it cannot complete (mirrors the skill blueprint field)
Model, Memory ScopeCapability and persistence. Memory defaults to none; use it only for genuine cross-run state (tracking an entity, learned preferences), and avoid it for research/freshness workflows where stale recall misleads — prefer a curated context file when the learning should stay human-visible.
Tools, SkillsCross-references to Integration Options entries and Skill IDs. Tools follow least privilege — only what the Responsibilities require.
Trigger Examples2-3 structured examples (context → user message → expected behavior → invocation) Build uses verbatim to construct <example> blocks in the agent’s description

When more than one agent is defined, this section captures the orchestration pattern (Supervisor / Pipeline / Parallel / Network), the coordinator, the Handoff Contracts table (data passed between agents), and the Aggregation Strategy.

Prerequisites lists platform setup, accounts, credentials, and plugin installs needed before the workflow can run. The Deployment Plan documents where each artifact lives and how it gets deployed — with a Packaging note explaining how artifacts ship together.

End of Layer 3. The only hard gate. After the model produces the spec, it is saved as a draft file you can open and read; say approve and it is marked approved. Build refuses an unapproved spec. Then move to Build.


These sections sit outside the three layers:

  • Evaluation Inputs — pointer to the Workflow Requirements’ Acceptance Criteria and Example Scenarios (not duplicated). Step 5 (Test) reads them from the Workflow Requirements directly.
  • Deferred to Build — explicit list of decisions intentionally left for Build (specific platform offering, exact model version, integration setup specifics)
  • Stakeholders (optional, organizational lens) — role swimlane and stakeholder details
  • Self-Test Summary — populated by the Design skill after running the Build Skill Needs Checklist; each item ✓ (passed) or ⚠️ (issue described inline). Lets you see what was verified before approving.

The three layers above are the conceptual structure of the Design Spec. In practice, the skill walks them across fourteen phases, in this chronological order:

  1. Load — Read the Workflow Requirements file from outputs/.
  2. Confirm understanding — Summarize the workflow and ask you to confirm.
  3. Architecture decisions — Layer 1: confirm platform (the one question), then extract tool integrations, trigger/schedule, and constraints from the Workflow Requirements and present a confirmation block.
  4. Autonomy — Assess where the whole workflow sits on the autonomy spectrum (Deterministic, Guided, Autonomous).
  5. Mechanism — Recommend a mechanism (Skill or Agent) with an involvement mode (Augmented or Automated).
  6. Safety & permissions — What the workflow may touch, what it may never do, and whether a write action is possible at all on your platform.
  7. Layer 1 confirmation — The whole architecture played back in plain English, with every term explained as it is confirmed.
  8. Classify each step — Layer 2: per-step autonomy level, AI building blocks, tools, human review gates.
  9. Skill discovery — Look for skills you already have before assuming anything must be built.
  10. Skill candidates — Steps tagged for skill creation with generation-ready detail.
  11. Agent configuration — Layer 3: when applicable, generate a platform-agnostic agent blueprint.
  12. Verify evaluation inputs — Confirm the acceptance criteria and example scenarios Test will grade against are complete.
  13. Write the draft spec — Write the complete design document as a draft.
  14. Approve — the draft spec is written to a file you read; say approve and it is marked approved. Build refuses anything else.

This step is facilitated by the design AI Workflow Framework skill. How you get it depends on your platform — see Install the Hands-on AI Plugin for installation.

How to start: Say “run the design skill” (or “design the workflow”) — works on every platform. With the plugin installed, Claude Code also accepts /handsonai:design, and Cowork lists it when you type /.

Platform compatibility: Claude (Chat, Cowork, Code) ✓  |  ChatGPT & Codex ✓  |  Gemini (Spark, Enterprise, CLI) ✓  |  M365 Copilot ✓  |  Cursor / Antigravity ✓

Start with this prompt:

Design the AI workflow from my Workflow Requirements.
Assess the autonomy level, recommend an orchestration mechanism, and map building blocks.

Upload or paste your Workflow Requirements file ([workflow-name]/requirements.md) from the Deconstruct step. The skill runs the three layers in order, with lightweight confirmation between each, and produces a Design Spec.

"Design the AI workflow from my Workflow Requirements"
→ Reads the most recent Workflow Requirements, runs Design,
produces the Design Spec for approval
"Design the expense-reporting workflow"
→ Reads outputs/expense-reporting-requirements.md, recommends
an orchestration mechanism, and generates the spec

The Design Spec is organized into the three layers above, plus cross-layer sections. The spec opens with YAML frontmatter so Build can summarize it in one read.

Before Layer 1: Source, then Value & Measurement — carried forward from the Workflow Requirements so the spec says what the workflow is for without a second file.

Layer 1 — Architecture sections: Execution Pattern, Architecture Decisions (with Packaging), Autonomy Spectrum Summary, Safety & Permissions (including Constraint Conformance), Integration Options (with Source URLs), Model Recommendation (with per-platform mapping).

Layer 2 — Decomposition sections: Step-by-Step Decomposition (or Capability Domain Mapping for goal-driven), Orchestrator Prompt Outline (when mechanism is Skill), Data Readiness Summary, Recommended Implementation Order.

Layer 3 — Component Blueprint sections: Skill Candidates (12 fields each), Agent Configuration (14 fields each; mandatory for goal-driven), Multi-Agent Configuration (when applicable), Prerequisites, Deployment Plan.

Cross-layer sections: Evaluation Inputs (pointer), Deferred to Build, Stakeholders (optional), Self-Test Summary (verification results).

Value & Measurement is copied forward from the Workflow Requirements unchanged — the objective, the outcome, the metric, today’s number, and the target. Design does not re-open those questions; it carries them so that whoever reads the spec on its own can still see what the build is for and how it will be judged. If today’s number was never established, the spec says so in those words rather than quietly dropping the row.

Constraint Conformance is the section that makes the protections from Deconstruct real. Design walks each one and records where the architecture satisfies it, in a small table with one of three states:

  • Satisfied — the design meets the constraint, and the row names how: which gate, which permission, which boundary.
  • Accepted — the constraint cannot be met as stated, and someone with the authority to say so has accepted the difference. The row names who accepted it.
  • Open — nothing in the design meets it yet.

An Open row is not a failure of the spec; it is the spec being honest. But it does not survive into a build: the design is not finished while one is outstanding, and Build reads them. A constraint that everyone assumed was handled, and that no document ever claimed was handled, is the one that fails in production.

Requirements written before this section existed simply have nothing to reconcile. Design notices that, fills in what it can infer from the architecture, and asks about the rest rather than declaring the workflow constraint-free by default.

Build reads both the Design Spec and the Workflow Requirements together — the Design Spec is the architectural blueprint; the Workflow Requirements remains the canonical source for per-step content, acceptance criteria, and human gates.


The Design Spec is structured for machine consumption — by the Build skill, by other AI agents, or by engineers building tooling around the spec.

ElementFormat
FrontmatterYAML. Required fields: workflow, requirements_file, spec_version, approved, definition_type, mechanism, involvement, platform, platform_mode, packaging, counts
Stable IDsS1, S2, … for skills; A1, A2, … for agents; C1, C2, … for context items (in the Workflow Requirements); E1, E2, … for example scenarios (in the Workflow Requirements)
Canonical vocabularyAutonomy: Human / Deterministic / Guided / Autonomous. Mechanism: Skill / Agent. Packaging: Plugin / Standalone Skill / Workspace Agent / Loose Files. Build Output: New skill: SN / Use existing: [name] / Extend existing: [name] / New agent: AN / Inline prompt → Workflow Requirements Step N / Handled by orchestrator / MCP server: [name] / Human (no artifact)
Section orderingLayer 1 → Layer 2 → Layer 3 → Cross-layer. Section names within each layer are fixed (consumers can locate any section by name).
Self-Test SummaryEach Build Skill Needs Checklist item marked ✓ or ⚠️ with inline description. Lets consumers see what was verified.

Consuming the spec standalone (without the Build skill)

Section titled “Consuming the spec standalone (without the Build skill)”

A capable agent (Claude Code, Codex, ChatGPT with skill support, Gemini CLI) can build the artifacts from the Design Spec alone, given access to:

  1. The Workflow Requirements file — referenced from frontmatter requirements_file:
  2. The agentskills.io specification — the open standard skill format (agentskills.io/specification)
  3. The target platform’s agent format docs — Claude Code’s sub-agents docs, Codex’s subagents docs, or equivalent

For best results, run Build via the framework’s Build skill — it auto-resolves these via the platform registry and handles deployment.

The Design Spec is portable across platforms because every major platform (Claude Code, Claude.ai, Cowork, Codex, ChatGPT, Gemini CLI) now follows the agentskills.io standard for skills. Three practical workflows:

  1. Same platform throughout (most common) — Design on Claude Code → Build on Claude Code. Platform in the spec matches Build’s runtime platform.
  2. Design once, Build elsewhere — Skill Candidates and Agent Configuration carry over unchanged because they’re agentskills.io-compatible. Update Architecture Decisions (Platform, Packaging) and Deployment Plan when moving. Example: design on Claude Code at home, build on Codex at work.
  3. Plan for a different target platform — Design on Claude Code targeting ChatGPT Workspace Agents (or any other platform). Build resolves the platform-specific translation at generation time.

agentskills.io is the portability layer. Every major platform converges on the standard, so Skill Candidates from one platform’s Design Spec build cleanly on another with minimal rewriting.