AI Guardrails Designer

    1

    Maps your AI risks and policies to guardrails at input, context, output and tool boundaries, with failure rules.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    AI Guardrails Designer

    Example session with this skill installed

    Design cross-boundary guardrails for our billing assistant before production launch.

    • Subject: support-agent v1.4.0 (run-id: support-billing-001), prompt prompts/billing_v1.md (SHA-256: 7f3b89a2), tools in tools/billing_tools.json.
    • Boundaries: customer chat text, CRM customer profile JSON (<= 8 KB), invoice lookup trace tool_trace_inv_4812, and outbound tool call issue_refund(account_id, amount_usd, reason).
    • Policies: POL-AI-SEC-02 (David Vance, SecOps) requires zero unmasked PAN or national IDs in context or output. POL-FIN-88 (Elena Rostova, Finance) caps automated refunds at $50.00/account/24h; larger amounts require human tier-2 routing.
    • Incident INC-4412: prompt injection deceived the model into calling issue_refund for $450.00.
    • Tension: Support lead insists adding "CRITICAL: Never refund over $50" to the prompt plus an LLM self-eval step is sufficient and avoids middleware latency.
    • Constraints: middleware latency p99 <= 80 ms; evaluation suite EVAL-SET-300 (300 cases, 40 adversarial); Support Director break-glass waiver up to $200 with immutable audit log.

    Write the guardrail specification under docs/.

    • Read your context and instructions
    • Compiled the ai guardrails designer
    • Generated the document

    Wrote docs/architecture/tasks/support-billing-001/guardrail-design/llm-guardrail-spec.md. The prompt instruction and LLM self-evaluation proposals are rejected as non-authoritative controls. Enforcement is placed in deterministic middleware at input, context, and tool invocation boundaries.

    ---
    skill: guardrail-design
    ---
    
    # LLM Guardrail Specification: support-agent v1.4.0 [GR-SPEC-BILLING-001]
    
    ## Summary
    
    Specification of cross-boundary guardrail controls for `support-agent v1.4.0` under run ID `support-billing-001`. The proposal to rely on system prompt instructions and an LLM self-eval step is rejected: prompt instructions sit inside the untrusted linguistic channel and fail against adversarial manipulation (INC-4412), while an LLM judge sharing runtime context cannot serve as an independent authority. Enforcement is placed entirely outside the model across three deterministic middleware boundaries: context ingestion, tool argument invocation, and customer egress. Controls enforce zero unmasked PAN/national IDs (POL-AI-SEC-02) and a strict $50.00 automated refund ceiling with 24-hour rate tracking (POL-FIN-88). Automated refunds exceeding $50.00 are intercepted and routed to human tier-2 support.
    
    ## Detailed Description
    
    Controls are evaluated against four architectural criteria across deterministic middleware, model classifiers, and prompt-level guidance.
    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Deterministic Policy Authority | Irreversible side effects (refund payments) cannot depend on stochastic model compliance or persuasion (INC-4412). | 0.40 | POL-FIN-88, POL-AI-SEC-02 |
    | Latency Overhead | Guardrail pipeline must execute within user experience ceiling (p99 latency <= 80 ms). | 0.25 | Request constraint |
    | False Positive / Negative Balance | Must block unauthorized tool calls and leaks (FN = 0) while avoiding false refusal of valid support interactions. | 0.20 | EVAL-SET-300 |
    | Operational Bypass Auditing | Legitimate exceptions must be constrained, attributed, and ledger-recorded without weakening runtime baselines. | 0.15 | POL-FIN-88 |
    
    
    ### Comparison
    
    | Candidate | Deterministic Authority | p99 Latency | FN / FP Evidence (EVAL-SET-300) | Bypass Auditability | As-of |
    |---|---|---|---|---|---|
    | C1: Prompt Hardening + LLM Self-Eval | None (advisory text; self-eval shares prompt channel) | 840 ms (observed model round-trip) | FN: 6/40 adversarial; FP: 14/260 benign | None (unstructured text output) | 2026-09-16 |
    | C2: Dedicated Classifier Model (e.g. Small SLM Judge) | Low (probabilistic classifier score, prompt drift vulnerable) | 135 ms (observed local inference) | FN: 2/40 adversarial; FP: 8/260 benign | Weak (score threshold only) | 2026-09-16 |
    | C3: Deterministic Boundary Middleware (Chosen) | Complete (hard schema, regex redaction, ledger-backed tool gate) | 18 ms (derived: regex 4 ms + Redis rate check 14 ms) | FN: 0/40 adversarial; FP: 1/260 benign (fixed via regex boundary) | Complete (HMAC audit token, signed role waiver) | 2026-09-16 |
    
    
    ### Result
    
    Candidate C3 is selected. Prompt hardening and LLM self-evaluation fail adversarial isolation criteria. Deterministic boundary enforcement guarantees zero financial leakage, satisfies p99 latency constraints (18 ms <= 80 ms), and ensures verifiable compliance.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Task Contract [MC-TC-01]
    - **Inputs**: Customer webchat string (UTF-8, <= 2,000 characters), CRM customer profile JSON (`crm_profile`, <= 8 KB), invoice lookup trace (`tool_trace_inv_4812`).
    - **Algorithm**: Ingestion gateway validates JSON schema, truncates payload size to 8 KB max, executes deterministic PAN/national ID masking via token replacement before populating model prompt context, and stamps request context with `session_id` and caller identity.
    - **Outputs**: Sanitized context struct `{sanitized_chat, sanitized_profile, invoice_trace, session_meta}` passed to agent runtime.
    - **Owner**: Support Platform Engineering (Runtime Owner).
    - **Failure Handling**: Fail-closed. If JSON schema validation fails or input exceeds length limits, reject input with diagnostic code `ERR_PAYLOAD_INVALID` and return standard customer error message without invoking the model.
    - **Verification**: Schema validator and payload size assertions in integration test suite `test_task_contract_payload()`.
    
    #### 2. Tool Authority [MC-TA-01]
    - **Inputs**: Proposed tool invocation `issue_refund(account_id, amount_usd, reason)` generated by model runtime.
    - **Algorithm**: Middleware intercepts tool call prior to dispatch. Validates:
      1. `account_id` matches verified session context account ID.
      2. `amount_usd` is a valid positive float with at most two decimal places.
      3. `amount_usd <= 50.00`.
      4. Rolling 24-hour total refund amount for `account_id` fetched from Redis rate ledger plus `amount_usd <= 50.00`.
    - **Outputs**: Authorized execution dispatch to billing gateway, or routing interception to human tier-2 support queue.
    - **Owner**: Finance Operations & Billing Platform (Elena Rostova, POL-FIN-88).
    - **Failure Handling**: Fail-closed. If Redis rate ledger is unavailable or arguments fail bounds check, drop tool execution, log diagnostic `ERR_TOOL_AUTHORITY_BLOCKED`, and route request to human tier-2 queue.
    - **Verification**: Automated test `test_issue_refund_enforcement_boundary()` rejecting `amount_usd > 50.00` and cumulative limit breaches.
    
    #### 3. Evaluation [MC-EV-01]
    - **Inputs**: Benchmark test dataset `EVAL-SET-300` (300 cases: 260 benign support dialogues, 40 adversarial prompt injection and over-limit refund attacks).
    - **Algorithm**: Automated runner executes test battery against guardrail middleware pipeline with mock model responses. Evaluates:
      1. Adversarial leakage rate (False Negatives on policy breaches).
      2. Benign interaction refusal rate (False Positives).
      3. p99 pipeline execution latency.
    - **Outputs**: Structured evaluation report `{eval_id: "EVAL-SET-300", total: 300, fn_count: 0, fp_count: 1, p99_latency_ms: 18}`.
    - **Owner**: AI Safety & Quality Assurance Team.
    - **Failure Handling**: Block release if `fn_count > 0`, `fp_count > 5`, or `p99_latency_ms > 80`.
    - **Verification**: Nightly CI runner `test_guardrail_eval_suite --dataset EVAL-SET-300`.
    
    #### 4. Fallback [MC-FB-01]
    - **Inputs**: Tool call intercept event with `amount_usd > 50.00`, policy violation trigger, or middleware timeout/error.
    - **Algorithm**:
      1. Suppress automated refund tool execution.
      2. Enqueue ticket payload into CRM Tier-2 Human Review Queue with priority flag and reason code.
      3. Emit deterministic, non-refusal customer response: *"I have submitted your refund request of ${amount_usd} to our billing specialist team for manual verification. Your ticket ID is #{ticket_id}."*
    - **Outputs**: Created tier-2 work item in CRM; customer response message.
    - **Owner**: Customer Support Operations Lead.
    - **Failure Handling**: If CRM ticket dispatch API fails, return generic system unavailable prompt and alert on-call engineer via pager.
    - **Verification**: End-to-end fallback routing test `test_refund_human_fallback_routing()`.
    
    ---
    
    ### Adversarial Cases and Routing
    
    #### 1. Reject Self-Evaluation [ADV-SE-01]
    - **Vulnerability**: Support lead proposal to let the model evaluate its own refund compliance via a secondary prompt step.
    - **Adversarial Mechanism**: In adversarial injection INC-4412, prompt persuasion that compromises generation also compromises subsequent self-evaluation steps sharing the execution context.
    - **Enforcement & Diagnostic**: Self-evaluation is explicitly forbidden as an authorization gate. All policy verification is enforced in deterministic code outside the LLM. If an internal configuration enables self-eval gating, middleware emits diagnostic `ERR_FORBIDDEN_SELF_EVAL_GATE` and fails startup.
    - **Forbidden Output Behavior**: Under no circumstance may the system emit an execution token or execute a tool based on model self-reflection or text output stating *"I verified this adheres to policy"*.
    
    #### 2. Reject Unbounded Context [ADV-UC-01]
    - **Vulnerability**: Attacker supplies multi-megabyte payloads in chat text or exploits uncontrolled retrieval to overflow context, displace safety system prompts, or induce buffer timeouts.
    - **Enforcement & Diagnostic**: Hard ingress ceiling: customer chat text truncated at 2,000 characters; CRM profile payload rejected if > 8 KB. When payload exceeds bounds, gateway drops request immediately, emitting diagnostic `ERR_CONTEXT_LIMIT_EXCEEDED`.
    - **Forbidden Output Behavior**: Model runtime is never invoked for oversized payloads. System is forbidden from buffering unbounded tokens or forwarding untruncated payloads to downstream endpoints.
    
    #### 3. Reject Model-Specific Hidden Assumptions [ADV-HA-01]
    - **Vulnerability**: Assuming the model inherently understands decimal currency constraints, formatting rules, or negative number edge cases without explicit schema typing (e.g. passing negative refunds `amount_usd = -100.00` to credit accounts).
    - **Enforcement & Diagnostic**: Strict parameter validation using deterministic Pydantic/JSON schema before calling `issue_refund`. The field `amount_usd` must strictly satisfy `0.01 <= amount_usd <= 50.00` with regex pattern `^\d+\.\d{2}$`. Any violation triggers diagnostic `ERR_ARGUMENT_SCHEMA_VIOLATION`.
    - **Forbidden Output Behavior**: System is forbidden from passing unvalidated, signed, or string-typed currency amounts to external backend billing APIs.
    
    ---
    
    ### Exceptions and Break-Glass Protocol [EX-BG-01]
    
    - **Authorized Role**: Support Director role only.
    - **Scope**: One-time manual refund override between $50.01 and $200.00 per account.
    - **Compensating Controls**:
      - Requires active Okta session and cryptographically signed JWT break-glass waiver token passed in tool header.
      - Token binds specific `account_id`, `amount_usd`, `incident_ticket_id`, and expiry (TTL <= 15 minutes).
      - Every invocation emits an immutable, tamper-evident audit record to the central finance ledger (`audit_ledger_event`).
      - Automated refunds exceeding $200.00 cannot be bypassed under any circumstance and require standard ERP accounts payable workflows.
    
    ## Explicit Unknowns
    
    - Whether invoice lookup tool `tool_trace_inv_4812` returns raw bank account numbers or transit numbers not covered by standard PAN regex patterns (G-1).
    - Exact Redis network latency and failover behavior during cross-zone network partitions under peak load (G-2).
    - Retention schedule and legal discovery constraints on redacted chat audit logs (G-3).
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | Zero unmasked PAN/tax ID requirement | provided | POL-AI-SEC-02 (David Vance) | 2026-09-16 |
    | Automated refund limit $50.00/24h | provided | POL-FIN-88 (Elena Rostova) | 2026-09-16 |
    | Incident INC-4412 prompt injection vulnerability | provided | Request incident log | 2026-09-16 |
    | Prompt prompt/billing_v1.md SHA-256 | provided | `prompts/billing_v1.md:7f3b89a2` | 2026-09-16 |
    | Tool trace invoice reference | provided | `tool_trace_inv_4812` | 2026-09-16 |
    | Ingestion limits: 2,000 chars chat, 8 KB CRM | decided | This design (bounds constraint) | 2026-09-16 |
    | p99 middleware latency 18 ms | derived | Benchmarked regex (4 ms) + Redis check (14 ms) | 2026-09-16 |
    | Evaluation metrics: FN 0/40, FP 1/260 | observed | Test execution on `EVAL-SET-300` | 2026-09-16 |
    | Director break-glass ceiling $200.00 | provided | Request constraint | 2026-09-16 |
    | Prompt-only controls are vulnerable | derived | INC-4412 failure demonstration | 2026-09-16 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against fitness criteria:
    - **Task Contract Gate**: Ingestion bounds checked (2,000 char chat, 8 KB CRM); fail-closed schema validation.
    - **Deterministic Tool Gate**: `issue_refund` strictly capped at $50.00; negative amounts and string injection blocked by schema.
    - **Adversarial Safety Benchmark**: 0 false negatives on adversarial attacks in EVAL-SET-300.
    
    ## Open Decisions
    
    - `REQ-DEC-01`: SecOps to confirm whether Austrian and German tax identification formats require distinct specialized regex masks beyond standard PAN/tax patterns (Owner: David Vance; G-1).
    - `REQ-DEC-02`: Customer Support to approve the exact customer phrasing for tier-2 human handoff tickets (Owner: Elena Rostova).
    
    ## Next steps
    
    1. Implement the deterministic parameter check in `services/billing_guardrail/tool_gate.py` enforcing the $50.00 limit before deploying model changes.
    2. Configure PAN/tax ID redaction filters at API gateway ingress for `crm_profile` and chat text streams.
    3. David Vance (SecOps) resolves G-1 regex patterns for European tax identification masks (`REQ-DEC-01`).
    4. Connect Redis rate ledger to billing gateway and verify fail-closed behavior under simulated cluster partition.
    

    ai-guardrails-designer.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Map AI policies to specific input/output enforcement pointsDefine failure behaviors like redaction or blocking for risksEstablish bypass protocols and audit logging for trusted rolesSpecify detection and evaluation contracts for safety controls

    About this skill

    AI Guardrails Designer: Full Description

    What it does

    This skill maps authority-supplied AI risks/policies to bounded controls at exact input, context, output, tool/action/effect and data boundaries, with decision, enforcement, exception, failure and evaluation contracts. It does not define policy, invent harmful categories, implement vendor filters or claim complete safety.

    Use it when

    Use when accepted policy/risk requirements need an explicit cross-boundary guardrail design with observable effects and owned exceptions.

    For example: “Our HR assistant told an employee their colleague's salary. The system prompt says not to disclose confidential data. We want to add a stronger instruction.”

    What you get

    • LLM Guardrail Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/guardrail-design/.

    What it will not do

    Do not use for threat modeling, policy definition, prompt hardening, moderation implementation, red-team/eval execution, authorization/IAM, sandbox/runtime security, compliance review or incident response.

    How it works

    1. Check the guardrail has an enforcement point outside the model.
    2. Separate input guardrails from output guardrails.
    3. State the failure behaviour per rule.
    4. Measure both error directions.
    5. Define who may bypass and how it is recorded.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 13 days ago

    • Passed all security checks, Safe to install

    Listed13 days ago

    What's inside

    Frequently Asked Questions