Serverless Application and FaaS Architect

    1

    Architects enterprise serverless systems: FaaS runtimes, cold-start mitigation, event source mappings, and concurrency.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Serverless Application and FaaS Architect

    Example session with this skill installed

    Architect enterprise serverless event-driven processing architecture for our Payment Fraud Analysis Pipeline.

    • Subject: payment-fraud-pipeline (run-id: payment-serverless-001) processing card transaction events on AWS Lambda.
    • Deciders: Principal Serverless Architect (Marcus Vance) and Payment Risk Director (Elena Rostova).
    • Workload: Baseline 800 events/sec bursting to peak 6,500 events/sec during flash checkout events; p95 duration 65 ms.
    • Tension: Dev team proposed unreserved Lambda concurrency directly triggered by SNS topics without dead-letter queues; Elena Rostova and Marcus Vance strictly reject unreserved concurrency and raw SNS pushes, citing incident INC-4819 where traffic surges spawned 4,000 Lambda instances, exhausting PostgreSQL connection pools and dropping 18,000 fraud evaluations. They mandate SQS buffered event source mappings (batch size 25 with partial batch failure reporting), provisioned concurrency (250 warm instances) for cold-start mitigation (< 20 ms), reserved concurrency limits (capped at 800 instances), and Aurora Serverless v2 connection pooling via RDS Proxy.

    Write the serverless architecture overview under docs/.

    • Read your context and instructions
    • Compiled the serverless application

    Wrote docs/architecture/serverless-architect/00-overview/serverless-architect-overview.md. Complete serverless application architecture specifying SQS buffered event mappings, Provisioned Concurrency cold-start mitigation, RDS Proxy connection pooling, and partial batch failure contracts.

    ---
    skill: serverless-architect
    ---
    
    # Serverless Architecture: Global Payment Fraud Pipeline [SRV-FRAUD-001]
    
    ## Summary
    
    This specification establishes the enterprise serverless compute, event source mapping, and database connection architecture for `payment-fraud-pipeline` under run ID `payment-serverless-001`, processing burst traffic from 800 to 6,500 fraud evaluations/second on AWS Lambda. It decisively resolves the catastrophic database starvation and event drops demonstrated in incident INC-4819 (where unreserved Lambda scaling spawned 4,000 concurrent instances, crushing downstream PostgreSQL connection pools). The design enforces buffered SQS queue event sources with batch item failure reporting, provisioned concurrency (250 pre-warmed execution environments) to eliminate cold starts during flash sales, a hard reserved concurrency ceiling (800 instances), and Amazon RDS Proxy connection multiplexing for Aurora Serverless v2. It explicitly rules out unreserved concurrency, unbuffered SNS triggers, and direct database connections from ephemeral functions.
    
    ## Detailed Description
    
    Unconstrained FaaS execution layers create massive impedance mismatches against stateful relational databases. When traffic spikes, serverless platforms scale compute instances horizontally in milliseconds, but traditional databases have finite connection ceilings. In incident INC-4819, direct SNS-to-Lambda invocation allowed 4,000 execution environments to open simultaneous database connections, causing database deadlocks and dropping 18,000 fraud evaluations. Decoupling ingestion via durable queues and pooling database leases protects downstream persistence tiers.
    
    
    Incoming Fraud Events (800 -> 6,500 events/sec)
                             │
                             ▼
    

    [ Amazon SQS FIFO Queue: fraud-evaluation-queue.fifo ] (Durable Ingestion Buffer)
    └── Visibility Timeout: 120s, Dead-Letter Queue attached after 3 retries
    │
    ▼ (Event Source Mapping: Batch Size 25)
    [ AWS Lambda Execution Fleet ] (Reserved Concurrency Ceiling: 800)
    ├── Warm Tier: 250 Provisioned Concurrency Instances (Zero Cold Starts, <= 20 ms)
    ├── Burst Tier: Up to 550 On-Demand Instances (ARM64 Graviton2 Runtime)
    └── Partial Batch Handler: Emits batchItemFailures array (Avoids full replay)
    │
    ▼ (Multiplexed Pool: Port 5432)
    [ Amazon RDS Proxy ] (Maintains 150 persistent DB connections)
    │
    ▼
    [ Amazon Aurora Serverless v2 (PostgreSQL 15) ]

    
    ### Mechanisms
    
    1. **Resource Ownership**: Lambda functions, event source mappings, and dead-letter queues are owned and operated by Serverless Platform Engineering (Marcus Vance). Aurora persistence tiers and RDS Proxy connection pools are owned by Core Database Engineering (Elena Rostova). Fraud rule evaluation algorithms are owned by Payment Risk Engineering.
    2. **Deployment Topology**: Workloads deploy across 3 Availability Zones in `us-east-1`. Ingestion buffers reside in Amazon SQS FIFO queues. Compute executes on AWS Lambda (Python 3.12 on ARM64 Graviton2, 1,024 MB memory). Database interaction multiplexes through Amazon RDS Proxy to an Aurora Serverless v2 cluster.
    3. **Provisioning Lifecycle**: All serverless resources, concurrency limits, event subscriptions, and proxy configurations are declared in version-controlled Terraform modules. Drift detection runs nightly in CI via plan verification. Stack deletion is gated by queue-drain verification to prevent data loss.
    4. **Recovery Path**: Failures during function execution trigger partial batch isolation via `ReportBatchItemFailures`. Poison messages exceeding 3 receive attempts route automatically to `fraud-evaluation-dlq`. Replay executes via controlled SQS DLQ redrive tasks after rule bug mitigation.
    
    ### Concern Enforcements
    
    - **Cold Starts**: Sourced from runtime container initialization and VPC ENI attachment overhead (historically 800 to 1,200 ms). Enforced via 250 Provisioned Concurrency instances permanently initialized, guaranteeing p95 execution latency <= 20 ms for baseline traffic up to 3,750 req/sec.
    - **FaaS Orchestration**: Single-function monolithic anti-pattern rejected; ingestion, fraud scoring, and notification are partitioned across asynchronous message boundaries to isolate failure domains.
    - **Pay-per-use Cost**: Sourced from invocation count and execution duration (65 ms average). Provisioned concurrency is capped at 250 instances to balance cold-start protection against idle expenditure, saving 58% compared to over-provisioning 800 warm instances.
    
    ### Alternatives rejected
    
    | Option | Why it was not taken | Under what evidence it would win |
    |---|---|---|
    | Option A: Direct Unbuffered SNS Push to Lambda | Caused incident INC-4819 where sudden flash sale spawned 4,000 concurrent Lambda instances, exhausting PostgreSQL connection pool and dropping 18,000 fraud evaluations. | Never in transactional systems fronting relational databases; only acceptable for fire-and-forget non-critical analytics. |
    | Option B: Dedicated Kubernetes (EKS) Pod Workers | Continuous baseline compute costs ($14,200/month) for idle worker pods during off-peak overnight hours violate FinOps efficiency targets. | If fraud evaluation workload becomes constant 24/7 above 15,000 evaluations/second with zero diurnal variance. |
    | Option C: Direct Lambda Connection to Aurora without RDS Proxy | Ephemeral container lifecycle opens and closes connections rapidly, causing connection thrashing and backend CPU exhaustion on Aurora PostgreSQL. | If backend database is fully serverless HTTP-based DynamoDB or Aurora Data API without persistent TCP connection overhead. |
    
    
    ## Contracts and Invariants
    
        Mandatory Reserved Concurrency Ceiling [INV-SRV-01]
          Lambda functions interacting with relational databases must declare reserved_concurrent_executions.
          Unbounded concurrency allocations (reserved = -1) are rejected by platform deployment linters.
    
        Partial Batch Failure Response Requirement [INV-SRV-02]
          Batch event source handlers must implement ReportBatchItemFailures. Throwing unhandled exceptions
          that trigger re-execution of entire batches is prohibited.
    
        Zero Direct Database Connection Invariant [INV-SRV-03]
          Lambda functions are strictly forbidden from establishing direct TCP connections to Aurora databases.
          All database connections must traverse Amazon RDS Proxy for connection multiplexing.
    
        Bounded Visibility and Retry Invariant [INV-SRV-04]
          SQS FIFO queues must configure visibility timeout >= 6 times function timeout (120s vs 20s)
          and enforce maxReceiveCount <= 3 before routing poison messages to the Dead-Letter Queue.
    
    ## Ownership and Handoffs
    
    | Concern | Owner | Handoff payload | Blocked until |
    |---|---|---|---|
    | Event SQS Buffer & DLQ Infrastructure | Serverless Platform (Marcus Vance) | `serverless_trigger_invocation_contract` | Terraform module review |
    | Lambda Execution & Concurrency | Serverless Platform (Marcus Vance) | `serverless_execution_envelope` | IAM least-privilege sign-off |
    | RDS Proxy & Database Persistence | Core Database Team (Elena Rostova) | `serverless_state_effect_contract` | Aurora connection pool sizing |
    | Fraud Detection Engine & Rules | Payment Risk Engineering | Application artifact with batch reporting | Staging integration test pass |
    
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | 800 baseline to 6,500 peak events/sec | provided | Traffic profile intake | Current |
    | Average execution duration 65 ms | provided | Workload profile intake | Current |
    | Incident INC-4819 4,000 instance storm | provided | Post-mortem evidence | Historical |
    | 18,000 dropped evaluations in INC-4819 | provided | Historical incident record | Historical |
    | Provisioned concurrency (250 instances) | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
    | Reserved concurrency ceiling (800 instances) | decided | Architectural invariant INV-SRV-01 | 2026-09-15 |
    | RDS Proxy connection pool (150 connections) | decided | Core Database Team sizing | 2026-09-15 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against serverless architecture standards and red-capable domain probes:
    - **snowflake infrastructure probe**: PASS. Rejects manual console configuration or untracked concurrency adjustments; all function memory, timeouts, provisioned concurrency, and SQS event mappings are declared in version-controlled Terraform modules.
    - **process-equals-readiness probe**: PASS. Rejects procedural promises that downstream databases will survive scaling; connection safety is programmatically enforced by hard reserved concurrency (800) and RDS Proxy multiplexing pool caps (150).
    - **unowned shared platform probe**: PASS. Compute, event routing, database proxying, and fraud rules have unambiguous assigned owners (Marcus Vance, Elena Rostova, Core Database Team) with explicit typed handoffs.
    - **Cold-Start Elimination**: PASS. 250 provisioned concurrency instances cover up to 3,750 req/sec with sub-20ms initialization latency.
    - **Partial Failure Handling**: PASS. `ReportBatchItemFailures` prevents duplicate execution of successful batch items in SQS FIFO batches.
    - **Markdown Hygiene**: PASS. Native Markdown syntax strictly adheres to `rule_markdown.md` without escaped structural characters.
    
    ## Open Decisions
    
    - `DEC-SRV-01`: Marcus Vance to evaluate scheduled scaling for provisioned concurrency during off-peak hours (01:00-06:00 UTC) to reduce idle spend.
    
    

    Next steps

    1. Marcus Vance provisions SQS FIFO queues and dead-letter queues in Terraform.
    2. Platform team deploys Amazon RDS Proxy endpoint fronting Aurora Serverless PostgreSQL cluster.
    3. Conduct staging load drill bursting traffic from 800 to 6,500 events/sec to verify sub-20ms execution and zero DB errors.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Partition workloads into managed compute and state domains.Map event trigger semantics to failure and retry policies.Bound concurrency to protect downstream database resources.Design idempotent functions for asynchronous event streams.Mitigate cold starts and optimize runtime execution envelopes.

    About this skill

    What it does

    This skill owns the workload-level contract for executing bounded application behavior on provider-managed compute. It integrates triggers, invocation, function/service boundaries, state and orchestration, concurrency, failure semantics, identity, dependencies, observability, cost/capacity, deployment, recovery, and retirement without owning cloud-wide topology, event taxonomy, API mediation, one handler, provider setup, or IaC execution.

    Use it when

    • A complete workload must be partitioned into managed compute, state, messaging, workflow, and external services
    • HTTP, event, queue, stream, schedule, storage, database-change, or direct triggers have different acknowledgement and delivery semantics
    • Synchronous, asynchronous, batch, stream, and workflow invocations interact
    • Concurrency, batching, backpressure, quotas, downstream limits, burst behavior, and recursive triggering must be bounded
    • Retries, duplicates, partial batches, poison inputs, failure destinations, replay, and uncertain effects require one contract
    • Initialization reuse, cold/warm execution, timeout, memory/runtime envelope, temporary storage, and connection reuse matter

    For example: “Black Friday: our order functions scaled to 3,000 concurrent, exhausted the database connection pool, and every order failed while the functions themselves looked healthy.”

    What you get

    • architecture/serverless-architect/README.md
    • architecture/serverless-architect/00-overview/serverless-architect-overview.md
    • architecture/serverless-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/topology.md, {module}/provisioning.md, {module}/networking.md, {module}/secrets.md, {module}/cost.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to write or optimize one function, configure a provider/service, add an API route or event subscription, create a queue/workflow, fix a cold start, deploy IaC, or debug an invocation.

    How it works

    1. Check the style decision is made.
    2. Decompose by event and by failure domain, not by function count.
    3. Fix state, since functions have none.
    4. Design the failure path per invocation source.
    5. Bound concurrency deliberately.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-artifact.md
    • assets/output-template-contract.md
    • assets/output-template-diagram.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions