- Home
- Skills
- Agents & Orchestration
- Multi-Agent Workflow Designer
Multi-Agent Workflow Designer
Designs a bounded multi-agent workflow: typed handoffs, graph limits, merge authority, failure and resume rules.
Free
Works with the AI tools you already use
Multi-Agent Workflow Designer
Example session with this skill installed
Design the workflow for our contract review assistant.
Today it spawns a researcher, a summariser and a critic that all share one context. Runs sometimes take 20 minutes, and last week the critic approved a summary that contained a clause that is not in the contract.
- Input: one contract PDF (at most 80 pages) and the customer's playbook (JSON, 40 clause types, each with an acceptable position).
- Output wanted: a risk report listing each clause, its type, the playbook position it breaks, and the page it is on.
- Tools: pdf_text(page) returns the text of one page; playbook_lookup(clause_type) returns the position. No write tools, no external calls.
- Budget: at most 60 model calls and 5 minutes per contract. A lawyer signs off every report before it is sent.
- Owner: the legal-ops team. Runtime: our existing Python orchestrator; no framework change.
Write the blueprint under docs/.
- Read your context and instructions
- Compiled the multi-agent workflow designer
- Generated the document
Wrote docs/architecture/tasks/contract-review-001/agent-workflow-design/multi-agent-workflow-blueprint.md. Short version: you do not need three agents. The job is a bounded pipeline with two model-controlled steps and a deterministic check that makes an invented clause impossible to report. It fits your budget at 48 model calls in the worst case.
Contract Review Workflow Blueprint [WF-CR-1]
Summary
The current researcher, summariser and critic share one context and differ only in prompt, so the critic re-reads the summariser's errors instead of the contract. That is how an invented clause was approved. This blueprint replaces them with four steps: extract, classify, verify, assemble. Only extract and classify call a model. A clause reaches the report only if its text is found on the cited page by a deterministic check. The graph has no cycles, at most 48 model calls, and a 5-minute hard stop that yields a report marked incomplete, never a silent failure.
Terminal Outcomes
| Outcome | When | Goes to |
|---|---|---|
ready_for_review | every page window and clause type processed; every reported clause verified | lawyer sign-off |
incomplete | time or call budget reached, or a tool failed after its retry | lawyer, with the unprocessed list |
rejected_input | not a PDF, more than 80 pages, or playbook JSON invalid | legal-ops |
The workflow never sends a report. Sending happens only after the lawyer signs off.
Workflow Graph
| Node | Type | Input | Output | Model calls |
|---|---|---|---|---|
| N1 extract | model step, per 10-page window | pdf_text of pages p..p+9 | {clause_id, text, page}[] | ≤ 8 |
| N2 classify | model step, per clause type found | clauses of one type + playbook_lookup(type) | {clause_id, type, breaks_position: bool, position_ref}[] | ≤ 40 |
| N3 verify | deterministic | clause text + pdf_text(page) | {clause_id, found_on_page: bool} | 0 |
| N4 assemble | deterministic | verified clauses + N2 results | risk report | 0 |
Worst case: 8 + 40 = 48 model calls, leaving 12 for retries inside the 60-call budget.
Edge Contracts
N1 → N2: JSON array, schema-validated. page must be inside the window; text must be non-empty. A malformed item is dropped and logged, never repaired by a model.
- N2 → N4: every
clause_idmust exist in N1 output; unknown IDs are rejected at the edge.
N1 → N3: every clause is checked against pdf_text(page) after whitespace normalisation. found_on_page: false removes the clause from the report and lists it under "rejected extractions".
N3 → N4: the report includes a clause only if found_on_page is true. This is the rule that stops last week's failure.
Graph Bounds
- Depth 2 (window → type); no node spawns another node.
- Fan-out: at most 8 windows, at most 40 clause types.
- Cycles: none. No re-review loop.
- Budget: 60 model calls and 5 minutes per contract, enforced by the orchestrator, not by prompts. Reaching either ends the run as
incomplete.
Merge and Approval Authority
- N4 is deterministic code; no model merges results.
- No node approves its own output: N3 checks N1, and the lawyer approves the report.
- The "critic" role is removed. A model asked to approve a predecessor that shares its context is not an independent check.
State, Retry and Resume
- Tools are read-only, so a repeated call has no side effect.
- Each window and each clause type is checkpointed when it completes. A resumed run skips completed work.
- One retry per failed model or tool call, only while budget remains. A second failure marks that window or type unprocessed.
Evaluation
- N3 guarantees no reported clause is absent from the contract.
- It does not catch missed clauses. Measuring misses needs contracts labelled by legal-ops (handoff H-3).
Handoffs
| ID | To | Needs |
|---|---|---|
| H-1 | prompt owner | N1 and N2 prompts that return the schemas above |
| H-2 | agent-architect owner | only if a later version needs an agent that chooses its own next action |
| H-3 | legal-ops | a labelled contract set to measure missed clauses |
Explicit Unknowns
- Which model runs N1 and N2, and its latency: the 5-minute limit is enforced, but not proven to fit 48 calls.
- How
pdf_texthandles scanned pages without a text layer. - Clauses that match more than one playbook type.
Traceability
| Claim | Classification | Source |
|---|---|---|
| 80-page limit, 40 clause types, tools, budget, lawyer sign-off | provided | request |
| 10-page windows, ≤ 48 calls | derived | 80 pages ÷ 10, plus 40 types |
| Shared context caused the invented clause | derived | request: one context, critic approved an absent clause |
| Deterministic verify and merge | decided | this design |
Verification
No validator was supplied, so no command was run. Self-checks against the request:
- call ceiling 48 ≤ 60;
- every edge has a schema and a rule for malformed input;
- no node certifies its own output.
Next steps
- Legal-ops confirms the three terminal outcomes and the "incomplete" report format.
- Prompt owner writes the N1 and N2 prompts against the edge schemas (H-1).
- Build N3 first: it is the cheapest step and closes the invented-clause failure on its own.
multi-agent-workflow-designer.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
About this skill
What it does
This skill defines one authority-bounded workflow in which deterministic, model-controlled, tool/service and human steps coordinate through typed state and handoffs. It does not invent agents, grant autonomy, select frameworks, implement runtime code or prove useful outcomes.
Use it when
Use when accepted AI-agent/runtime capabilities need a concrete workflow/control graph with explicit steps, transitions, contracts, effects, failures and human intervention.
For example: “Our contract review assistant spawns a researcher, a summariser and a critic. It sometimes runs for twenty minutes, and last week the critic approved a summary that had invented a clause.”
What you get
- Multi-Agent Workflow Blueprint
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/agent-workflow-design/.
What it will not do
Do not use for integrated autonomous-agent architecture, deterministic workflow implementation, prompt/tool/memory/model design alone, operating coding agents or generic task delegation.
How it works
- Check more than one decision loop is warranted.
- Assign each node a task class, not a persona.
- Fix the contract on every edge.
- Bound the graph.
- Decide who merges and who may not self-certify.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
4 people have installed this skill.
Trust & safety
Security scanned
Verified clean 14 days ago
- Free to download with an account