- Home
- Skills
- Agents & Orchestration
- ML Serving Architecture Selector
Works with the AI tools you already use
ML Serving Architecture Selector
Choose and justify batch, real-time, streaming, edge, managed, Kubernetes, or serverless model serving with an evidence-tagged ADR.
$19
ML Serving Architecture Selector
Example session with this skill installed
We need to serve a 2GB LLM for internal document Q&A. We have 500 users, p99 latency needs to be under 2s, and our team has zero Kubernetes experience. Batch or real-time?
- Read your context and instructions
- Compiled the ml serving architecture
- Generated the UI component
ARTIFACT: ML-SERVING-ADR
DECISION: Managed real-time endpoint; avoid self-hosted Kubernetes.
BINDING CONSTRAINTS
- [Verified] p99 latency under 2 seconds eliminates batch for the interactive path.
- [Verified] Zero Kubernetes experience eliminates self-hosted Kubernetes.
COST: requests/day × compute-seconds × current unit price; CALIBRATE request rate and vendor price.
STEELMAN: A 2GB model may create cold-start or memory constraints that require direct measurement.
REVISIT: Re-open the choice when measured concurrency, latency, cost, or team capacity crosses an approved threshold.
ml-serving-architecture-selector.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Choose a serving pattern your team can actually operate
ML serving decisions are often made from technology preference instead of workload constraints. This skill turns latency, throughput, freshness, cost shape, team capacity, rollout needs, and compliance requirements into a structured architecture decision.
What it produces
- A binding-constraint register tagged Verified, Inferred, or Assumed
- An elimination pass for structurally unsuitable patterns
- A six-axis comparison matrix for the surviving options
- Cost formulas using supplied values and CALIBRATE markers where values are missing
- An ML-SERVING-ADR with the decision, strongest counterarguments, and revisit triggers
Patterns covered
Batch, real-time, streaming, edge, managed platforms, Kubernetes-based serving, and serverless inference. The skill compares architecture patterns; it does not deploy or configure any platform.
Best for
- Teams deciding whether online inference is actually necessary
- Architects comparing managed serving with self-operated infrastructure
- ML teams documenting why an option was eliminated
- Reviewers who need explicit trade-offs and conditions that reopen the decision
Boundaries
It does not query live prices, benchmark latency, provision infrastructure, prove compliance, or guarantee cost or performance. Provider-specific claims must be verified against current authoritative documentation.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 1 month ago
- Passed all security checks, Safe to install