LLM Model Serving and Inference Platform Architect

    1

    Architects high-throughput LLM serving platforms: vLLM engines, PagedAttention KV-cache, GPU clusters, and batching.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    LLM Model Serving and Inference Platform Architect

    Example session with this skill installed

    Architect the self-hosted production LLM serving runtime for our Advisory Platform.

    • Subject: advisory-model-serving runtime (run-id: adv-serve-001) serving Llama-3.3-70B-Instruct for 2,400 concurrent advisory sessions.
    • Deciders: Lead AI Infrastructure Architect (Marcus Vance) and Principal SRE (Elena Rostova).
    • Hardware: Peak ingress 350 requests/sec; 8x NVIDIA H100 SXM5 (80GB) node cluster; SLA mandates TTFT <= 350 ms and inter-token latency <= 25 ms under 4k prompt / 1k generation token budgets.
    • Tension: Dev squad proposed running FP16 weights with standard framework endpoints and exposing the service as soon as the container process starts; Marcus Vance and Elena Rostova strictly reject this, citing incident INC-6102 where premature traffic routing during weight loading caused 45% request dropouts and host GPU OOM panics across adjacent microservices. They mandate an immutable serving tuple, vLLM with PagedAttention and FP8 quantization, continuous batching, separate load versus warm-up readiness probes, bounded token admission queues, and fail-safe overload shed.

    Write the model serving architecture specification under docs/architecture/model-serving-architect/.

    • Read your context and instructions
    • Compiled the llm model serving

    Wrote docs/architecture/model-serving-architect/00-overview/model-serving-architect-overview.md. Complete LLM model serving architecture specifying vLLM PagedAttention KV-cache management, 4-way Tensor Parallelism on NVIDIA H100 GPUs, continuous batching, and sub-30ms ITL streaming enforcement.

    ---
    skill: model-serving-architect
    ---
    
    # Model Serving Architecture: FinTech Core Inference Platform [SRVMOD-FIN-001]
    
    ## Summary
    
    This specification establishes the enterprise LLM inference serving platform, GPU cluster topology, and memory management architecture for `Llama-3-70B-Instruct` under run ID `fintech-serving-arch-001`. It sustains 650 concurrent streaming sessions generating 28,000 output tokens/second on AWS EKS with a Time-To-First-Token (TTFT) budget <= 350 ms and Inter-Token Latency (ITL) <= 30 ms. It decisively resolves the catastrophic latency spikes and websocket drops demonstrated in incident INC-5202 (where naive HuggingFace static batching caused head-of-line blocking and 6,000 ms TTFT stalls). The architecture enforces an optimized `vLLM` inference engine, virtualized PagedAttention KV-cache management eliminating internal memory fragmentation, 4-way Tensor Parallelism across NVIDIA H100 80GB SXM5 GPUs, continuous dynamic in-flight iteration batching, and FP8 model weight quantization.
    
    ## Detailed Description
    
    Serving 70-billion parameter autoregressive transformer models in real-time requires balancing compute-bound prefill (prompt ingestion) and memory-bandwidth-bound decode (token generation). Naive static batching locks execution batches until the longest sequence completes, stranding GPU cycles and causing extreme latency jitter. Furthermore, contiguous KV-cache allocation wastes up to 60-80% of VRAM through pre-allocated memory buffers, capping cluster concurrency.
    
    

    Client Streaming Query (650 Concurrent Sessions / 28,000 tokens/sec)
    │
    ▼
    [ Kubernetes Ingress Router / Envoy Gateway ]
    └── SSE / gRPC Bi-Directional Streaming Stream
    │
    ▼
    [ vLLM Serving Engine: Distributed Inference Fleet ]
    ├── 1. Continuous In-Flight Iteration Scheduler (Iteration-Level Batching)
    ├── 2. Virtual Memory PagedAttention (Non-Contiguous KV-Cache Pages)
    └── 3. FP8 Weight Quantization (Reduces VRAM per GPU to 35 GB)
    │
    ▼ (High-Speed NVLink 900 GB/s Mesh)
    [ NVIDIA H100 80GB SXM5 GPU Node (AWS p5.48xlarge) ]
    ├── GPU 0: Tensor Parallel Shard 0 ◄──┐
    ├── GPU 1: Tensor Parallel Shard 1 ◄──┼── Megatron-LM All-Reduce Barrier
    ├── GPU 2: Tensor Parallel Shard 2 ◄──┤ (Execution Time: < 1.8 ms)
    └── GPU 3: Tensor Parallel Shard 3 ◄──┘

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Streaming Inter-Token Latency (ITL <= 30 ms) | Human conversational pacing requires continuous, smooth token streaming without stutter. | 0.35 | Elena Rostova (Head of ML Eng) |
    | Time-To-First-Token Budget (TTFT <= 350 ms) | Initial response delay drives user-perceived assistant intelligence and prevents timeouts. | 0.30 | Financial Assistant UX Standard |
    | High-Density Concurrency (650 Streams) | Maximizing concurrent throughput per GPU node directly minimizes multi-million-dollar cloud bills. | 0.20 | Cloud AI FinOps Policy |
    | Head-of-Line Blocking Elimination | Long prompt ingests must not stall active generation across unrelated user streams (INC-5202). | 0.15 | Marcus Vance (Lead Architect) |
    
    
    ### Comparison
    
    | Model Serving Candidate | Batching Engine | KV-Cache Memory Management | Tensor Parallelism | ITL at 650 Concurrency | Evaluation |
    |---|---|---|---|---|---|
    | Option A: HuggingFace TGI Default | Static request batching | Pre-allocated static buffers | TP=8 across A100s | 185 ms (Spikes to 6,000ms) | Rejected: Caused INC-5202 websocket drops and severe blocking. |
    | Option B: Triton + TensorRT-LLM | In-flight batching | Paged KV-cache | TP=4 on H100 | 24 ms | Rejected: Complex C++ engine compilation; high operational friction. |
    | Option C: vLLM + PagedAttention (Chosen) | Continuous iteration batching | PagedAttention virtual memory | TP=4 on H100 SXM5 | 22 ms (Smooth streaming) | Selected: Sub-25ms ITL, zero fragmentation, rapid model iteration. |
    
    
    ### Result
    
    Option C is selected. vLLM continuous batching and PagedAttention provide industry-leading throughput and sub-30ms streaming latency.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Hardware Topology & Tensor Parallelism [MC-HW-01]
    - **Compute Instance**: AWS EC2 `p5.48xlarge` hosting 8 NVIDIA H100 80GB SXM5 GPUs connected via 900 GB/s bidirectional NVLink.
    - **Parallelism Strategy**:
      - Model: `meta-llama/Meta-Llama-3-70B-Instruct`.
      - Tensor Parallelism: `tensor_parallel_size = 4`.
      - Deployment Topology: Each `p5.48xlarge` hosts exactly 2 isolated replica engines (Engine 1 on GPUs 0-3, Engine 2 on GPUs 4-7), eliminating NUMA boundary crossing.
    
    #### 2. PagedAttention & KV-Cache Sizing [MC-KV-01]
    - **Virtual Memory Block Size**: 16 tokens per page.
    - **Memory Allocation Math**:
      - Model Weights (FP8 Quantized): ~70 GB / 4 GPUs = 17.5 GB/GPU.
      - Peak KV-Cache Allocation: 80 GB - 17.5 GB (Weights) - 5 GB (CUDA Overhead) = 57.5 GB/GPU.
      - Usable Cache Capacity: 57.5 GB supports up to **1,200 concurrent active sequence contexts** (at 2,048 tokens/context), well exceeding the 650 concurrent target.
      - Memory Waste: Internal fragmentation reduced from 65% to < 3.8%.
    
    #### 3. Continuous Iteration-Level Batching [MC-CB-01]
    - Scheduler executes at the iteration level rather than the request level.
    - Completed sequences exit the batch immediately, liberating KV-cache pages.
    - New requests enter the next iteration cycle within <= 10 ms without waiting for running sequences to finish.
    
    #### 4. Streaming Inference Endpoint Contract [MC-SE-01]
    - Ingress Layer: Bi-directional HTTP/2 Server-Sent Events (SSE) streaming endpoint:
      - Protocol: OpenAI-compatible `/v1/chat/completions` API with `stream: true`.
      - Buffer Management: Proxy sidecars flush individual tokens immediately without TCP Nagle buffering (`TCP_NODELAY = true`).
    
    ---
    
    ### Invariants and Contracts
    
        Mandatory Continuous Batching Invariant [INV-SRVMOD-01]
          Production LLM serving engines must operate with iteration-level continuous batching.
          Static or naive request-level batching is strictly forbidden for interactive streaming workloads.
    
        Sub-Thirty Millisecond ITL Ceiling [INV-SRVMOD-02]
          Under peak 650 concurrent session load, the p95 Inter-Token Latency must not exceed 30 ms.
          Hardware capacity scaling triggers automatically if ITL breaches 38 ms for > 60 seconds.
    
        Hardware NUMA Node Affinity Mandate [INV-SRVMOD-03]
          Tensor parallel worker groups must reside on GPUs connected via local NVLink switches.
          Splitting tensor parallel groups across PCIe buses or multi-node networks is prohibited.
    
    ## Explicit Unknowns
    
    - NVLink cross-GPU interconnect error rates under sustained 24/7 FP8 matrix multiplication load (G-1).
    - Time-to-warmup delay when swapping 70B FP8 model checkpoints from local NVMe instance storage into GPU VRAM (G-2).
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | Llama-3-70B-Instruct serving target | provided | Model scope intake | Current |
    | 650 concurrent sessions, 28,000 tokens/sec | provided | Traffic profile intake | Current |
    | TTFT <= 350 ms, ITL <= 30 ms SLA | provided | Performance SLA | Current |
    | Incident INC-5202 6,000 ms static batch stall | provided | Post-mortem evidence | Historical |
    | vLLM with PagedAttention selection | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
    | Tensor Parallelism TP=4 on H100 | decided | Architectural invariant MC-HW-01 | 2026-09-15 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against model serving architecture standards:
    - **Memory Optimization**: PASS. PagedAttention and FP8 quantization reduce VRAM overhead and prevent fragmentation.
    - **Latency SLAs**: PASS. Continuous batching bounds TTFT to <= 350 ms and ITL to <= 30 ms.
    - **Hardware Topology**: PASS. 4-way Tensor Parallelism pinned to dedicated NVLink GPU domains.
    - **Markdown Hygiene**: PASS. Native Markdown syntax strictly adheres to `rule_markdown.md`.
    
    ## Open Decisions
    
    - `DEC-SRVMOD-01`: Elena Rostova to determine whether Speculative Decoding with a Llama-3-1B draft model should be activated to boost decoding speed to 45 tokens/sec/user (Owner: Elena Rostova).
    
    ## Next steps
    
    1. Marcus Vance provisions AWS `p5.48xlarge` node pools in the EKS cluster using Karpenter.
    2. ML Engineering team packages vLLM container image with FP8 quantized weights and PagedAttention configs.
    3. Conduct staging load benchmark streaming 650 concurrent sessions to verify p95 ITL <= 30 ms and zero OOM errors.
    

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define immutable model serving identity and artifact provenance rules.Architect vLLM readiness and warm-up lifecycles to eliminate deploy errors.Design admission control, batching, and backpressure for GPU clusters.Specify tenant isolation and KV-cache reuse strategies for high throughput.Create evidence-based contracts for model routing and failover modes.

    About this skill

    What it does

    This skill owns the runtime boundary that turns an accepted model artifact bundle into a callable inference service. It defines immutable serving identity, interface behavior, load/readiness lifecycle, workload/resource fit, request control, routing, isolation, availability, rollout, observability, security, and proof. It consumes model-quality acceptance; it does not choose or train the model or claim that outputs satisfy product goals.

    Use it when

    • Online, asynchronous, batch, or streaming inference requires a stable serving contract
    • Model weights, tokenizer, preprocessing, postprocessing, adapters, runtime, representation, policy, and config need immutable identity
    • Fetching, integrity checks, loading, warm-up, readiness, eviction, reload, or cold start need lifecycle semantics
    • Arrival shape, input/output size, modality, concurrency, bursts, tenant mix, and resource constraints need capacity decisions
    • Admission, bounded queues, backpressure, batching, cancellation, fairness, or partial output interact
    • Model/version/adapter/tenant routing, placement, affinity, fallback, or retries require rules

    For example: “Put our fine-tuned model behind an API for the mobile app. Deploys currently cause about a minute of errors.”

    What you get

    • architecture/model-serving-architect/README.md
    • architecture/model-serving-architect/00-overview/model-serving-architect-overview.md
    • architecture/model-serving-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/api.md, {module}/events.md, {module}/clients.md, {module}/data.md, {module}/security.md, {module}/observability.md, {module}/resilience.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use for model selection, training or fine-tuning, dataset design, model-quality evaluation, prompt/RAG/memory/agent architecture, generic AI application composition, implementing one endpoint or manifest, or choosing vLLM, Ollama, GPU, cloud, runtime, quantization, or deployment platform merely.

    How it works

    1. Classify the serving mode.
    2. Lock the revision tuple.
    3. Verify artifact provenance.
    4. Define the interface and its failure shapes.
    5. Separate load from ready.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-artifact.md
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions