- Home
- Skills
- Agents & Orchestration
- AI Agent Architect
AI Agent Architect
Designs an autonomous AI agent's runtime: authority, tool and side-effect controls, stop rules, delegation, oversight.
$12
Works with the AI tools you already use
AI Agent Architect
Example session with this skill installed
Design the runtime decision loop for our refund agent (RefundOps).
Support engineers spend 14 minutes per ticket manually checking dispute ledgers and fraud scores. We want RefundOps to autonomously process eligible low-risk refunds.
- Subject: RefundOps v1.2 evaluating disputed merchant transactions.
- Owner: Elena Rostova (VP Fintech Engineering). Decider: Risk Operations Guild.
- Trigger: Webhook dispute_opened(case_id, tx_id, user_id, amount_usd).
- Outcomes: refunded, escalated_risk, escalated_limit, rejected_ineligible, blocked_error.
- Tools: ledger_read(tx_id), fraud_score_lookup(user_id), issue_refund(tx_id, amount, reason, idempotency_key), escalate_to_human(case_id, reason).
- Envelope: Auto-refund <= $150/tx; daily customer cap $300; fraud score must be < 0.65.
- Bounds: Max 8 turns, 24k context tokens, 45-second wall clock, $0.20 budget per run.
- Tension: Support proposed prompt-only rules ('instruct LLM to respect $150 limit') and self-graded evaluation ('ask model if refund is compliant'). Risk insists tool proxy must hard-block > $150 and an independent ledger oracle must verify balance changes.
Write the architecture under docs/architecture/agent-architect/.
- Read your context and instructions
- Compiled the ai agent architect
- Generated the document
Wrote docs/architecture/agent-architect/00-overview/agent-architect-overview.md. Complete agent runtime architecture specifying authority limits, step and tool budgets, human oversight gates, and anti-pattern verification.
---
skill: agent-architect
---
# AI Agent Architecture: RefundOps Agent [AA-REFUND-001]
## Summary
This specification establishes the bounded autonomous runtime for RefundOps Agent, handling merchant disputes and consumer refund claims. The agent operates under hard programmatic authority limits: maximum $150 per automated transaction, daily user refund limit of $300, and mandatory human escalation for fraud risk scores >= 0.65 or claims above $150. It explicitly rules out autonomous policy overrides, self-adjudication, and direct ledger credit creation outside the guarded `issue_refund` tool interface.
## Detailed Description
RefundOps processes refund requests received via inbound support cases. The agent analyzes transaction ledger data, queries the fraud score oracle, evaluates eligibility against refund policy rules, and either executes a bounded refund or delegates the case to human risk analysts.
Incoming Case
│
▼
[ Runtime Envelope: 30 Tools / 120s ]
│
├─► ledger_read ──► Verify Transaction
├─► fraud_score_lookup ──► Check Risk Metric
│
├── Fraud Score >= 0.65 OR Amount > $150 ──► escalate_to_human
└── Fraud Score < 0.65 AND Amount <= $150 ──► issue_refund
### Alternatives rejected
| Option | Why it was not taken | Under what evidence it would win |
|---|---|---|
| Direct $300 Auto-Approval Ceiling | Fraud chargebacks increased by 18% in the preceding month; raising ceiling doubles exposed financial risk. | Reversal requires 60 consecutive days of fraud chargeback rates dropping below 1.0%. |
| Prompt-Only Guardrails | LLM system instructions cannot guarantee financial ceilings; prompt injection could bypass threshold. | Rejected permanently; financial thresholds must be deterministically enforced by tool wrappers. |
| Fully Manual Review | High ticket volume creates unacceptable resolution backlogs for trivial low-value claims. | Reversal if customer dispute volume drops below 50 tickets/day. |
## Contracts and Invariants
Financial Authority Ceiling [INV-REFUND-01]
The agent runtime wrapper strictly enforces that `issue_refund` cannot be invoked with `amount` > 150.00 USD, regardless of LLM reasoning or prompt injection. Violations immediately abort execution with error code ERR_AUTH_CEILING.
Daily Aggregate Customer Limit [INV-REFUND-02]
Total cumulative refunds issued to a single `user_id` within a rolling 24-hour window must not exceed 300.00 USD. If cumulative sum + requested amount > 300.00, execution delegates to `escalate_to_human`.
Mandatory Escalation on Risk Score [INV-REFUND-03]
Any case returning `fraud_score` >= 0.65 from `fraud_score_lookup` must trigger `escalate_to_human`. The agent is denied authority to call `issue_refund` when risk threshold is breached.
Execution Resource Bounding [INV-REFUND-04]
Runtime hard bounds: max 30 model tool iterations, 120 seconds wall-clock timeout. Upon exceeding either limit, the run aborts and falls back to a parked status in the human operator queue.
## Ownership and Handoffs
| Concern | Owner | Handoff payload | Blocked until |
|---|---|---|---|
| Financial Authority & Limits | Risk Operations Guild (Elena Rostova) | Authority ceiling rules & escalation threshold definitions | Approved by Elena Rostova |
| Tool Wrappers & Ledger Integration | Core Payments Engineering | OpenAPI schema for `ledger_read` and `issue_refund` | Wrapper integration test passes |
| Human Review Console | Customer Support Operations | UI case queue format for `escalate_to_human` | Queue webhook receiver deployed |
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Auto-approval max $150 | provided | Intake specification | Current |
| Daily customer limit $300 | provided | Intake specification | Current |
| 18% fraud chargeback increase | provided | Risk operations report | Stated in request |
| Max 30 tool calls / 120s budget | provided | Intake specification | Current |
| Deterministic wrapper enforcement | decided | Architectural decision AA-REFUND-001 | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against fitness anti-patterns:
- **Unbounded Autonomy Check**: PASS. Hard limits enforced: 30 tool calls, 120s timeout, $150 per transaction, $300 daily per user.
- **Self-Graded Evaluation Check**: PASS. Agent does not grade its own decisions; risk scores originate from independent `fraud_score_lookup` oracle.
- **Prompt-Only Control Check**: PASS. Ceilings enforced at tool boundary in backend wrapper code, not via LLM system prompt instructions.
## Open Decisions
None. All constraints and boundaries derived directly from intake request.
Next steps
- Review deterministic tool wrapper enforcement logic with Core Payments Engineering.
- Verify human review queue ingestion schema with Customer Support Operations lead.
- Establish weekly audit reconciliation between agent-issued refunds and ledger records.
ai-agent-architect.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the integrated runtime architecture for an AI-controlled loop that can select actions from observations rather than execute only a fixed, owner-authored sequence. It defines the agent boundary, goal and authority, observation–decision–action lifecycle, context and state, tool/effect controls, planning and stopping, delegation, human intervention, interruption/resume, failure containment, evaluation hooks and evidence.
Use it when
- An AI-controlled loop decides whether and how to continue after each observation
- A goal must be translated into bounded actions without silently expanding authority
- Context, working state and external memory influence decisions across turns or steps
- Tools have discovery, eligibility, input/output, permission and side-effect boundaries
- Plans may be created, revised, abandoned or escalated from observed results
- Termination, inability, uncertainty, interruption and human escalation need explicit outcomes
For example: “Our support tool should read a ticket, look up the customer's orders, and issue a refund under $50 without a human. Above $50 an agent escalates.”
What you get
- architecture/agent-architect/README.md
- architecture/agent-architect/00-overview/agent-architect-overview.md
- architecture/agent-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/api.md, {module}/events.md, {module}/clients.md, {module}/data.md, {module}/security.md, {module}/observability.md, {module}/resilience.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use for prompt wording alone, one function-calling schema, memory/RAG/model-serving design alone, deterministic workflow orchestration, implementing a named agent framework, or adding multiple agents merely.
How it works
- Confirm the loop is adaptive.
- Name the goal and every terminal outcome.
- Inventory action capabilities.
- Draw the loop once.
- Set the stop conditions.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 13 days ago
- Passed all security checks, Safe to install