- Home
- Skills
- APIs & Backend
- Workflow Orchestration Engine Architect
Workflow Orchestration Engine Architect
Architects workflow orchestration clusters: durable state engines, activity worker topologies, and cluster resilience.
$9
Works with the AI tools you already use
Workflow Orchestration Engine Architect
Example session with this skill installed
Architect enterprise Workflow Orchestration Engine cluster topology, state persistence, and worker fleet for Global Trade Finance.
- Subject: trade-finance-orchestrator-cluster v2.0 (run-id: trade-wforch-arch-001) orchestrating 90-day Letters of Credit and cross-border customs settlements.
- Deciders: Lead Orchestration Systems Architect (David O'Reilly) and Head of Trade Operations Infrastructure (Elena Rostova).
- Scope: Peak 8,500 active executions/sec; 140,000 concurrent long-running deal instances; sub-20ms state commit latency across AWS EKS.
- Tension: Squad deployed a standalone Temporal developer server with local SQLite persistence on a single EC2 instance; David O'Reilly and Elena Rostova strictly reject standalone instances, citing incident INC-4946 where disk corruption obliterated 420 active Letters of Credit, stranding $64M in shipping cargo for 8 days. They mandate an Enterprise Workflow Orchestration Architecture: multi-node Temporal cluster on AWS EKS, AWS Aurora PostgreSQL multi-AZ persistence with WAL archiving, dedicated activity worker autoscaling pools, and disaster recovery replication.
Write the workflow orchestration architecture overview under docs/.
- Read your context and instructions
- Compiled the workflow orchestration engine
- Generated the document
Wrote docs/architecture/workflow-orchestration-architect/00-overview/workflow-orchestration-architect-overview.md and docs/architecture/workflow-orchestration-architect/verification/fitness-self-check.md. Complete enterprise workflow orchestration specification establishing distributed Temporal cluster topologies, Aurora multi-AZ persistence, dedicated activity worker fleets, and cross-region disaster recovery.
skill: workflow-orchestration-architect
Workflow Orchestration Architecture: Global Trade Finance [WFORCH-TRADE-001]
Summary
This specification establishes the enterprise Workflow Orchestration Engine architecture, cluster deployment topology, state store persistence, and activity worker fleet scaling for trade-finance-orchestrator-cluster v2.0 under run ID trade-wforch-arch-001. It governs long-running Letters of Credit (LC), bills of lading verification, and cross-border customs settlement workflows spanning 90-day lifecycles across 140,000 concurrent active deal executions sustaining 8,500 peak state transitions/second. It decisively resolves the single-node disk corruption and data loss demonstrated in catastrophic incident INC-4946 (where running an un-clustered standalone developer server with SQLite storage on a single instance destroyed the active state of 420 international trade transactions, stranding $64M in shipping containers at maritime ports for 8 days). The architecture enforces a highly available, distributed Temporal.io cluster on multi-AZ AWS EKS, pairs with
AWS Aurora PostgreSQL 16 multi-AZ clustered persistence, establishes
partitioned activity worker autoscaling fleets, enforces
cross-region active-passive disaster recovery, and guarantees an internal state commit latency of
p99 <= 20 ms.
Detailed Description
Operating long-running financial workflows on single-node or un-clustered orchestration engines creates unacceptable single-point-of-failure risks. If the orchestrator's state store is corrupted or its compute host crashes, multi-million-dollar contractual commitments lose their execution state. A production-grade workflow orchestration architecture partitions the engine into independent, stateless service tiers (Frontend, History, Matching, Worker) backed by an enterprise-grade distributed relational database cluster. Stateless services scale independently based on ingress gRPC traffic and internal task matching queues, while the persistent state layer guarantees linearizable, ACID-compliant history event appends.
Public Partner Ingress & Bank APIs (8,500 state transitions/sec)
│
▼ (gRPC Ingress over mTLS)
[ Frontend Service Fleet (6 Replicas across 3 AZs) ]
├── 1. Ingress Rate Limiting & Auth Token Check
└── 2. Routes Workflow Signals & Queries to Matching Service
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
[ History Service ] [ Matching Service ] [ Internal Worker Fleet ]
(Shard Allocator) (Task Queue Sync) (System Maintenance)
│ │
▼ ▼ (Polls Activity Queues)
[ AWS Aurora PostgreSQL 16 ] [ Activity Worker Pods (KEDA HPA) ]
├── Multi-AZ Primary + Replica├── `trade-customs-worker-pool`
└── Sub-20ms ACID Commit Log └── `trade-wire-clear-worker-pool`
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Cluster High Availability & Zero Data Loss | An engine failure must never corrupt active Letters of Credit state (INC-4946). | 0.40 | David O'Reilly (Chief Orchestration Architect) |
| Long-Running Concurrency Scale (140,000 Deals) | System must retain and track 140,000 concurrent 90-day active trade workflows. | 0.30 | Elena Rostova (Head of Trade Operations) |
| State Transition Commit Latency (p99 <= 20 ms) | Core history service commits must execute in sub-20ms to prevent worker dispatch lag. | 0.15 | Core Trade Settlement SLA |
| Cross-Region Disaster Recovery (RPO=0, RTO < 2m) | Regional cloud outages must not freeze maritime shipping releases. | 0.15 | Maritime Trade Regulatory Policy |
Alternatives rejected
| Option | Why it was not taken | Under what evidence it would win |
|---|---|---|
| Standalone Developer Server (SQLite / H2) | Caused INC-4946 ($64M stranded cargo); single point of failure; zero high availability. | Local developer desktop testing with zero production traffic. |
| Custom Microservice Cron Pollers | Re-creates database race conditions and lacks event-sourced replay audit trails. | Simple stateless scheduled tasks running less than 1 minute. |
| Clustered Temporal on EKS + Aurora (Chosen) | Retains selection; enterprise multi-AZ HA, sub-20ms ACID state commits, proven scale. | Mission-critical enterprise platforms managing high-value multi-day financial contracts. |
Contracts and Invariants
Distributed Cluster Redundancy Invariant [INV-WFORCH-01]
All core orchestration cluster services (Frontend, History, Matching) must run with a minimum of 3 replicas
distributed across at least three distinct AWS Availability Zones. Single-instance deployments are strictly barred.
Zero Data Loss Storage Durability [INV-WFORCH-02]
The orchestrator persistence store must be configured with Multi-AZ synchronous replication and point-in-time
recovery (PITR). SQLite, ephemeral disks, or un-replicated databases are prohibited in production.
Activity Worker Queue Isolation Mandate [INV-WFORCH-03]
Activity worker pools must be isolated by functional domain onto dedicated Kubernetes node groups.
High-compute document generation workers must not share task queues with payment clearing workers.
Ownership and Handoffs
| Concern | Owner | Handoff payload | Blocked until |
|---|---|---|---|
| Orchestration Cluster Topology & EKS Infra | Chief Orchestration Architect (David O'Reilly) | temporal_cluster_helm_values | Multi-AZ EKS cluster ready |
| Trade Workflow Models & Queue Quotas | Head of Trade Operations (Elena Rostova) | trade_lifecycle_queue_matrix | Trade committee sign-off |
| Aurora PostgreSQL State Store Configuration | Lead Database Administrator | aurora_temporal_db_provisioning | Terraform staging release |
| Cross-Region DR Standby Cluster | SRE Disaster Recovery Squad | cross_region_temporal_replication | AWS us-west-2 VPC peering |
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 140,000 concurrent long-running trade deals | provided | Trade platform sizing profile | Current |
| 90-day Letters of Credit lifecycle | provided | Business scope intake | Current |
| Peak 8,500 state transitions/sec | provided | Volumetric performance intake | Current |
| Incident INC-4946 $64M stranded port cargo | provided | Historical forensic post-mortem | Historical |
| State commit latency budget p99 <= 20 ms | provided | Trade Settlement SLA | Current |
| Clustered Temporal + Aurora selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Multi-AZ 3-replica minimum invariant | decided | Architectural invariant INV-WFORCH-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against workflow orchestration architecture standards:
- Redundancy Rigor: PASS. 6 Frontend, 8 History, and 6 Matching replicas across 3 AZs guarantee zero SPOF.
- Persistence Safety: PASS. Aurora PostgreSQL Multi-AZ with PITR prevents repeat of INC-4946 disaster.
- Worker Isolation: PASS. Dedicated activity queues prevent resource contention between document and payment tasks.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-WFORCH-01: David O'Reilly to determine whether Temporal Server version 1.24 with advanced SQL visibility or Elasticsearch 8 visibility cluster is standardized for trade query filtering (Owner: David O'Reilly).
Next steps
- Platform Engineering deploys the production Temporal cluster via Helm on AWS EKS.
- Database team provisions the AWS Aurora PostgreSQL 16 cluster with 32 vCPU primary and read replica.
- Conduct staging disaster recovery drill simulating complete termination of the primary AWS AZ to confirm zero workflow state loss.
skill: workflow-orchestration-architect
Trade Finance Workflow Orchestration Platform — Fitness Self-Check [WFORCH-TRADE-FIT-001]
Summary
This fitness self-check evaluates the workflow orchestration platform architecture against three critical red-capable domain failure probes: anemic model, cross-context transaction, and duplicate language. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Anemic Model | Seed an activity worker implementation that bypasses domain models and executes arbitrary raw SQL updates directly against the orchestrator's internal history database. | Kubernetes network policy probe probe_unauthorized_internal_db_access verifying connection drop with diagnostic ERR_DIRECT_ORCHESTRATOR_DB_MUTATION_PROHIBITED. | pass | Confirms Calico/Cilium network firewall policies; does not evaluate root terminal access by cluster admins. |
| FIT-2: Cross-Context Transaction | Seed an activity task that attempts to open a distributed 2PC transaction locking both the trade finance database and the customs broker external database. | Activity execution policy linter probe_activity_distributed_lock_rejection verifying build rejection on 2PC library imports with diagnostic ERR_DISTRIBUTED_ACTIVITY_LOCK_PROHIBITED. | pass | Confirms activity codebase static analysis; does not inspect external third-party API transaction scopes. |
| FIT-3: Duplicate Language | Seed an activity worker queue definition that redefines core trade finance concepts (LetterOfCredit, BillOfLading) inconsistently with the canonical trade dictionary. | Activity schema dictionary validator probe_activity_vocabulary_drift verifying build rejection on mismatched Protobuf schemas with diagnostic ERR_ACTIVITY_VOCABULARY_DRIFT_DETECTED. | pass | Confirms Protobuf schema registry validation; does not inspect internal worker code comments. |
Residual Risk
- Cross-region active-passive synchronization lag (up to 1,500 ms) under WAN network fiber degradation. Accepted by Elena Rostova with read-only query routing on the secondary standby region during failover.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of direct orchestrator DB access | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of distributed activity 2PC locks | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of activity vocabulary drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates workflow orchestration fitness probes into automated CI deployment verification.
- Platform team configures Prometheus alerts monitoring Temporal History service state commit latencies and task queue backlog depths.
- Conduct quarterly disaster recovery drill validating automated failover to the secondary standby AWS region.
workflow-orchestration-engine-architect.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the application architecture for a coordinator that tracks and advances a long-running process across participants, failures, waits, restarts, deployments and operator intervention. It defines workflow identity and state, commands/events/activities, side-effect boundaries, durability, concurrency, timeout/cancellation/retry/idempotency, compensation, human tasks, version/replay compatibility, evidence and recovery while preserving domain and participant authority.
Use it when
- A process runs across multiple services/processes or over a duration where callers cannot hold one transaction/request open
- One coordinator must own execution identity, legal states, transitions, terminal outcomes and progress authority
- Participants need explicit command, event, activity, result, signal or query contracts
- Process state/history/checkpoints must survive worker or coordinator interruption
- Concurrent branches, joins, races, late results or duplicate starts need deterministic business outcomes
- Deadlines, timers, cancellation, retry exhaustion and manual recovery change workflow state
For example: “Customer onboarding involves six systems. When it fails halfway, someone works out what happened from logs and fixes it by hand. Last month two accounts were created twice.”
What you get
- architecture/workflow-orchestration-architect/README.md
- architecture/workflow-orchestration-architect/00-overview/workflow-orchestration-architect-overview.md
- architecture/workflow-orchestration-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/api.md, {module}/events.md, {module}/clients.md, {module}/data.md, {module}/security.md, {module}/observability.md, {module}/resilience.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use for AI-agent task planning, a simple cron job, data/batch pipeline design, BPMN diagrams, vendor workflow configuration, one event handler/API retry, or choosing an engine because words such as workflow, orchestration, saga, durable, scheduler, pipeline, approval, or compensation appear.
How it works
- Check central orchestration is warranted.
- Model the workflow as durable state, not as a call chain.
- Separate the deterministic workflow from its side effects.
- Define compensation for every step with an external effect.
- State timeouts, retries and human intervention points.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install