Works with the AI tools you already use

    CClaude CodeCCursorCCodex CLIGGitHub CopilotGGemini CLI+17 more

    ML Serving Architecture Selector

    1

    Choose and justify batch, real-time, streaming, edge, managed, Kubernetes, or serverless model serving with an evidence-tagged ADR.

    Secure checkout via Stripe

    0 installsSecurity scanned

    See it in action

    You say

    We need to serve a 2GB LLM for internal document Q&A. We have 500 users, p99 latency needs to be under 2s, and our team has zero Kubernetes experience. Batch or real-time?

    Your agent does

    ARTIFACT: ML-SERVING-ADR DECISION: Managed real-time endpoint; avoid self-hosted Kubernetes. BINDING CONSTRAINTS: - [Verified] p99 latency under 2 seconds eliminates batch for the interactive path. - [Verified] Zero Kubernetes experience eliminates self-hosted Kubernetes. COST: requests/day × compute-seconds × current unit price; CALIBRATE request rate and vendor price. STEELMAN: A 2GB model may create cold-start or memory constraints that require direct measurement. REVISIT: Re-open the choice when measured concurrency, latency, cost, or team capacity crosses an approved threshold.

    What you get

    Select between batch, real-time, streaming, or edge inference patterns.Generate defensible ADRs with explicit trade-offs and steelman arguments.Calculate infrastructure cost projections using real workload variables.Align architecture choices with a team's actual operational maturity level.

    About this skill

    Choose a serving pattern your team can actually operate

    ML serving decisions are often made from technology preference instead of workload constraints. This skill turns latency, throughput, freshness, cost shape, team capacity, rollout needs, and compliance requirements into a structured architecture decision.

    What it produces

    • A binding-constraint register tagged Verified, Inferred, or Assumed
    • An elimination pass for structurally unsuitable patterns
    • A six-axis comparison matrix for the surviving options
    • Cost formulas using supplied values and CALIBRATE markers where values are missing
    • An ML-SERVING-ADR with the decision, strongest counterarguments, and revisit triggers

    Patterns covered

    Batch, real-time, streaming, edge, managed platforms, Kubernetes-based serving, and serverless inference. The skill compares architecture patterns; it does not deploy or configure any platform.

    Best for

    • Teams deciding whether online inference is actually necessary
    • Architects comparing managed serving with self-operated infrastructure
    • ML teams documenting why an option was eliminated
    • Reviewers who need explicit trade-offs and conditions that reopen the decision

    Boundaries

    It does not query live prices, benchmark latency, provision infrastructure, prove compliance, or guarantee cost or performance. Provider-specific claims must be verified against current authoritative documentation.

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    Listedtoday

    Frequently Asked Questions