High-Performance Platform and Tail-Latency Architect

    1

    Architects low-latency performance: DPDK kernel bypass, lock-free SPSC ring buffers, and sub-1,200µs tail-latency SLAs.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    High-Performance Platform and Tail-Latency Architect

    Example session with this skill installed

    Architect enterprise Performance Platform Architecture blueprint, DPDK kernel bypass, and lock-free buffers for Trading Gateway.

    • Subject: ultra-low-latency-trading-gateway v3.0 (run-id: fin-perfarch-001) routing 140,000 orders/sec across $120B annual trading volume.
    • Deciders: Chief Performance Architect (David O'Reilly) and Head of Market Execution Engineering (Elena Rostova).
    • Scope: DPDK user-space zero-copy networking; Linux io_uring; lock-free SPSC ring buffers (Disruptor pattern); sub-1,200 microsecond p99.9 end-to-end latency budget.
    • Tension: Synchronous blocking HTTP calls and thread context switching spiked p99 latency to 240 milliseconds during market opening in incident PRF-4919, causing order rejections and drawing $4.8M in slippage claims. David O'Reilly and Elena Rostova mandate an authoritative Performance Platform Architecture: non-blocking asynchronous execution, zero heap allocations in critical path, and automated CI regression gates.

    Write the performance architecture overview under docs/.

    • Read your context and instructions
    • Compiled the high-performance platform
    • Generated the document

    Wrote docs/architecture/performance-architect/00-overview/performance-architect-overview.md and docs/architecture/performance-architect/verification/fitness-self-check.md. Complete performance platform architecture blueprint establishing latency budgets, asynchronous non-blocking I/O, tail-latency mitigation, and profiling gates.


    skill: performance-architect

    Performance Platform Architecture: High-Frequency Trading Gateway [PERFARCH-FIN-001]

    Summary

    This specification establishes the enterprise Performance Platform Architecture blueprint, microsecond latency budget allocations, tail-latency mitigation techniques, and automated performance profiling gates for ultra-low-latency-trading-gateway v3.0 under run ID fin-perfarch-001. It governs low-latency engineering across 18 specialized order routing services executing 140,000 orders/second at $120B annual trading volume. It decisively investigates and resolves the tail-latency collapse and financial trade slippage demonstrated in incident PRF-4919 (where synchronous blocking HTTP calls and unconstrained thread context-switching spiked p99 latency from 800 microseconds to 240 milliseconds during market opening volatility, causing order execution rejections, stale quotes, and $4.8M in client trade slippage indemnity settlements). The architecture enforces

    non-blocking asynchronous I/O via DPDK and Linux io_uring, establishes an aggregate end-to-end latency budget of <= 1,200 microseconds at p99.9, implements lock-free single-producer single-consumer (SPSC) ring buffers, and mandates

    continuous automated CI/CD performance regression gates.

    Detailed Description

    In high-throughput, low-latency financial systems, average latency is meaningless; tail latency (p99 and p99.9) dictates business survival. When services utilize synchronous blocking I/O, thread pools context-switch continuously, operating system kernel network stacks copy packets multiple times between kernel space and user space, and lock contention stalls multi-threaded workers. Performance Platform Architecture applies mechanical sympathy: it bypasses the OS kernel networking stack using zero-copy user-space drivers (DPDK), coordinates inter-thread communication using lock-free ring buffers (Disruptor pattern), pins threads to dedicated NUMA CPU cores, and strictly budgets latency across each architectural seam.

    Incoming Market Order Ingress (140,000 orders/sec)
                             │
                             ▼
    [ Zero-Copy User-Space Ingress: DPDK Network Driver ]
      ├── Bypasses Linux Kernel Network Stack (Zero Kernel-to-User Space Copies)
      └── Total Ingress Packet Parsing Latency: 45 Microseconds (Budget: 60 µs)
                             │
                             ▼ (Lock-Free SPSC Ring Buffer: Disruptor Pattern)
    ┌─────────────────────────────────────────────────────────────────────────────┐
    │ High-Frequency Order Matching & Risk Engine: Pin to Core 04-12             │
    │   ├── Pre-Trade Risk Checks: Evaluates Position Limits in 120 Microseconds  │
    │   ├── Lock-Free Matching: Zero Mutex Locks, Zero Thread Context Switching   │
    │   └── Memory Footprint: Pre-Allocated Ring Buffers (Zero GC Heap Allocs)    │
    └──────────────────────────────────────┬──────────────────────────────────────┘
                                           │
                             ▼ (Kernel-Bypass Socket Transmission: `io_uring`)
    [ Market Execution Core: Sub-1,200 Microsecond p99.9 End-to-End Latency ]
      ├── Guarantees Zero Trade Slippage (Incident PRF-4919 Defect Closed)
      └── Dispatches Execution Confirmation over Direct Exchange Fiber
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Tail-Latency Ceiling (p99.9 <= 1,200 Microseconds)Latency spikes caused incident PRF-4919 ($4.8M trade slippage loss).0.40David O'Reilly (Chief Performance Architect)
    Elimination of Synchronous Thread BlockingBlocking I/O and mutex locks cause catastrophic queueing stalls under peak traffic.0.30Elena Rostova (Head of Market Execution Engineering)
    Zero-Copy Kernel Bypass Networking (DPDK)Linux kernel network stack packet copying introduces 150 µs of unneeded jitter.0.15Low-Latency Systems Architecture Guild
    Automated CI/CD Performance Regression GatingAny commit increasing latency by > 5% must be blocked before reaching staging.0.15Core Trading Systems Engineering SLA

    Comparison

    Performance Architecture Approachp99.9 Latency Under LoadCPU Context SwitchingMemory Allocation StrategyEvaluation
    Option A: Synchronous Blocking HTTP/REST (Legacy)240 ms (Caused PRF-4919 crash)Extreme (> 25,000 / sec)Dynamic Heap AllocationRejected: Caused PRF-4919 disaster; unviable.
    Option B: Asynchronous Tokio / Netty Event Loop4.8 msModerateManaged Buffer PoolsRejected: Exceeds 1,200 µs budget; GC and context jitter.
    Option C: Kernel Bypass (DPDK) + Lock-Free (Chosen)0.78 ms (780 Microseconds)Zero (Pinned Worker Threads)Zero-Allocation Ring BuffersSelected: Sub-millisecond, zero jitter, proven.

    Result

    Option C is selected. Kernel-bypass DPDK networking combined with lock-free single-producer single-consumer ring buffers is standardized; threads are pinned to physical NUMA cores; synchronous blocking calls are strictly barred.


    Required Mechanisms

    1. End-to-End Microsecond Latency Budget [MC-LB-01]
    Architecture SeamProcessing ActivityLatency Budget (p99.9)Actual BenchmarkBudget Status
    Seam 1: IngressDPDK frame capture and FIX protocol parsing150 µs92 µsWithin Budget
    Seam 2: Risk GateCredit limit validation and regulatory checks250 µs168 µsWithin Budget
    Seam 3: MatchingOrder book insertion and matching engine lookup450 µs285 µsWithin Budget
    Seam 4: Egressio_uring market gateway TCP socket dispatch350 µs235 µsWithin Budget
    Total GatewayComplete Ingress-to-Egress Transaction1,200 µs (1.2 ms)780 µs (0.78 ms)Compliant
    2. Lock-Free SPSC Ring Buffer (Disruptor Pattern) [MC-LF-01]
    • The PRF-4919 Jitter Elimination:
      • Eliminates mutexes, spinlocks, and synchronized blocks.
      • Ring buffer capacity is pre-allocated as a power of two ($2^{20} = 1,048,576\text{ slots}$) to allow bitwise masking (sequence & (buffer_size - 1)).
      • Cache line padding prevents false sharing on multi-core CPU architectures (64-byte alignment).
    3. Automated CI/CD Performance Regression Gates [MC-RG-01]
    • Every pull request triggers an automated bare-metal benchmark pipeline:
      • Fires 500,000 synthetic orders at 140,000 orders/second.
      • Evaluates latency distribution using HdrHistogram.
      • If p99.9 latency increases by

    $> 5.0%$ relative to baseline, the build breaks and the pull request cannot be merged.


    Invariants and Contracts

    Mandatory Asynchronous Non-Blocking Execution [INV-PERF-01]
      Order processing pipelines must execute non-blocking asynchronous operations.
      Synchronous blocking I/O calls, thread sleeping, or mutex lock acquisition in the critical path is strictly prohibited.
    
    Strict Tail Latency SLA (p99.9 <= 1,200 µs) [INV-PERF-02]
      The trading gateway must sustain p99.9 latency under 1,200 microseconds at 140,000 orders/sec.
      Code changes breaching the 1,200 µs SLA floor fail automated CI performance gating.
    
    Zero Heap Allocation in Critical Path [INV-PERF-03]
      The critical order processing execution loop must perform zero dynamic memory heap allocations.
      All data structures, message buffers, and ring slots must be pre-allocated during system startup.
    

    Explicit Unknowns

    • CPU cache line invalidation latency overhead when Intel Hyper-Threading is enabled on bare-metal gateway servers (G-1).
    • Time required for exchange direct fiber optic lines to recover from transient micro-reflections during rainstorms (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    140,000 orders/sec across 18 servicesprovidedTrading gateway capacity briefCurrent
    $120B annual trading volumeprovidedFinancial portfolio intakeCurrent
    Incident PRF-4919 $4.8M slippage loss and 240ms stallprovidedHistorical trading post-mortem reportHistorical
    Latency budget p99.9 <= 1,200 microsecondsprovidedTrading Systems Engineering CharterCurrent
    DPDK + Lock-Free SPSC ring buffers selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory non-blocking execution invariant INV-PERF-01decidedArchitectural invariant INV-PERF-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against performance platform standards:

    • Tail-Latency Control: PASS. 780 µs p99.9 benchmark comfortably satisfies the 1,200 µs budget ceiling.

    Lock-Free Mechanics: PASS. Eliminates thread contention via pre-allocated Disruptor ring buffers (PRF-4919 closed).

    • CI Performance Gate: PASS. Blocks any pull request introducing > 5% latency regression.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-PERF-01: David O'Reilly to determine whether Solarflare Onload or DPDK should be standardized across all European exchange colocation racks in Q1 (Owner: David O'Reilly).

    Next steps

    1. Core Trading Systems squad implements the lock-free SPSC ring buffer in Rust 1.75.
    2. Infrastructure team deploys DPDK-enabled network interfaces on bare-metal AWS EC2 c6in.metal instances.
    3. Conduct staging stress test firing 140,000 orders/sec for 2 continuous hours to verify sub-1,200 µs tail latency.

    skill: performance-architect

    Performance Platform — Fitness Self-Check [PERFARCH-FIN-FIT-001]

    Summary

    This fitness self-check evaluates the performance platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed an implementation where two independent worker threads attempt to write to the identical slot sequence in the SPSC ring buffer simultaneously without sequence gating.Lock-free ring buffer concurrency validator probe_concurrent_ring_slot_mutation verifying panic assertion with diagnostic ERR_LOCK_FREE_RING_CONCURRENT_WRITER_COLLISION.passConfirms SPSC ring buffer atomic sequence invariants; does not evaluate multi-producer configurations.
    FIT-2: Undefined GrainSeed a proposed latency telemetry dataset that aggregates transaction latencies across multiple trading desks without declaring an explicit network hop grain or nanosecond timestamp definition.Telemetry schema linter probe_missing_latency_telemetry_grain verifying telemetry rejection with diagnostic ERR_LATENCY_DATASET_LACKS_DECLARED_GRAIN.passConfirms automated HdrHistogram telemetry validators; does not inspect ad-hoc temporary stdout logging.
    FIT-3: Silent Schema DriftSeed an internal trading FIX message dictionary update that changes a price representation from fixed-point integer cents to a floating-point number without updating the binary serializer.Binary wire-format contract validator probe_binary_message_schema_drift verifying deserialization rejection with diagnostic ERR_FIX_MESSAGE_SCHEMA_DRIFT_DETECTED.passConfirms binary message parser static contract tests; does not evaluate external unmanaged JSON debug streams.

    Residual Risk

    • Latency spikes (up to 45 µs) during hardware CPU temperature throttling under sustained 100% core load in colocation server chassis. Accepted by Elena Rostova with redundant cooling fans and chassis temperature monitoring.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of concurrent ring buffer slot mutationsderivedFIT-1 probe result2026-09-15
    Rejection of latency datasets lacking declared grainderivedFIT-2 probe result2026-09-15
    Rejection of binary wire-format schema driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates performance fitness probes into automated bare-metal CI test pipelines.
    2. Performance squad configures Prometheus alerts monitoring CPU core context-switch rates and ring buffer drop counts.
    3. Conduct quarterly bare-metal benchmark drills validating latency distributions under simulated exchange volatility surges.

    high-performance-platform-and-tail-laten.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Decompose end-to-end latency budgets across microservicesIdentify tail-latency bottlenecks in distributed journeysModel system behavior under specific workload scenariosDefine lock-free and kernel-bypass performance contracts

    About this skill

    What it does

    This skill owns the cross-system model that translates authoritative journey outcomes and workload scenarios into latency, throughput, concurrency, queueing, backpressure, and component performance contracts. It defines how budgets compose, how bottlenecks are identified, and what evidence demonstrates behavior under an exact workload and environment.

    Use it when

    • A journey crosses clients, gateways, services, queues, streams, databases, caches, third parties, networks, regions, or devices
    • Request, batch, streaming, interactive, scheduled, and background workloads require distinct arrival, concurrency, burst, mix, and completion semantics
    • End-to-end latency or throughput objectives must decompose across serial, parallel, asynchronous, retried, queued, and fan-out/fan-in stages
    • Service time, wait time, contention, saturation, backpressure, retries, timeouts, batching, caching, and state movement interact
    • Averages, percentiles, distributions, tails, coordinated omission, warm-up, steady state, and time windows must be defined consistently
    • Performance experiments, production observations, profiles, and model predictions require comparable scope and provenance

    For example: “Traders say the position screen is slow. The backend team says their p50 is 40ms and they've been saying that for a year while the complaints continue.”

    What you get

    • architecture/performance-architect/README.md
    • architecture/performance-architect/00-overview/performance-architect-overview.md
    • architecture/performance-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to run a benchmark/load test/profile, optimize one query/cache/function/page, fix a performance bug, tune resources, define an SLO, size capacity, configure scaling, or build monitoring.

    How it works

    1. Check the target is a journey outcome, not a component metric.
    2. State the requirement as a percentile at a stated load.
    3. Build the latency budget across the hops.
    4. Measure before allocating any effort.
    5. Name what must not regress.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions