- Home
- Skills
- Agents & Orchestration
- AI System Architect
AI System Architect
Designs end-to-end AI system architecture: model lifecycle, data pipelines, serving topology, evaluation, and governance.
$12
Works with the AI tools you already use
AI System Architect
Example session with this skill installed
Architect the end-to-end AI system for our Medical Imaging Diagnostic Assistant.
- Subject: MediScan AI Platform v1.0 analyzing chest X-ray DICOM images for pulmonary abnormalities (run-id: mediscan-arch-001).
- Deciders: Chief Medical Officer (Dr. Aris Thorne) and AI Systems Lead (Marcus Vance).
- Workload: Ingestion of 15,000 DICOM studies daily; inference latency p95 <= 2,500 ms; continuous learning loop.
- Infrastructure: Hybrid architecture with on-premise GPU cluster (4x NVIDIA H100 nodes) for HIPAA-compliant image analysis, and private cloud for audit metadata.
- Tension: Data science team wants to automatically retrain the production vision model on weekly inference inputs and auto-deploy new weights; Dr. Thorne strictly forbids autonomous clinical weight deployment, mandating human radiologist shadow evaluation (minimum 5,000 cases with F1 >= 0.94) and FDA-compliant sign-off before weight promotion.
- Governance: Mandatory model lineage tracking (data hash + model weights + eval report); deterministic fallback to radiologist priority queue upon model confidence < 0.85.
Write the AI system architecture overview under docs/.
- Read your context and instructions
- Compiled the ai system architect
- Generated the UI component
Wrote docs/architecture/ai-system-architect/00-overview/ai-system-architect-overview.md. Complete end-to-end AI system architecture establishing DICOM ingestion pipelines, on-premise GPU serving topology, shadow clinical evaluation gates, and human-in-the-loop abstention fallbacks.
---
skill: ai-system-architect
---
# AI System Architecture: MediScan Diagnostic Platform [AI-SYS-MED-001]
## Summary
This specification establishes the end-to-end architecture for the MediScan Diagnostic Platform v1.0 under run ID `mediscan-arch-001`. MediScan analyzes chest X-ray DICOM images across 15,000 daily patient studies for pulmonary abnormalities using an on-premise cluster of 4x NVIDIA H100 nodes within a 2,500 ms p95 inference budget. It decisively rejects autonomous clinical weight deployment in favor of an audited continuous learning loop: new model weights require shadow clinical evaluation across 5,000 historical studies achieving $F_1 \ge 0.94$ and formal Chief Medical Officer approval. The design enforces deterministic human radiologist routing whenever model classification confidence falls below 0.85.
## Detailed Description
Clinical AI systems operate under strict regulatory and patient-safety constraints. Unlike standard SaaS workloads, machine learning models exhibit non-deterministic stochastic behaviors, silent data drift, and sensitivity to distribution shifts in sensor hardware.
DICOM Study Ingress (15,000/day)
│
▼
[ De-identification & Normalization Gateway ]
│
▼
[ On-Premise GPU Inference Cluster (4x H100) ]
│
┌────────┴────────┐
▼ ▼
(Confidence ≥ 0.85) (Confidence < 0.85)
│ │
[ Diagnostic Assist ] [ Immediate Radiologist Priority Queue ]
│
▼
[ Clinical Feedback & Shadow Evaluation Store ] ──► (5k Shadow Gate F1 ≥ 0.94) ──► CMO Approval
### Alternatives rejected
| Option | Why it was not taken | Under what evidence it would win |
|---|---|---|
| Autonomous Continuous Retraining & Deployment | Unchecked model weight updates risk catastrophic forgetting and severe patient misdiagnosis in production. | Never. Clinical healthcare regulations (FDA/HIPAA) require deterministic validation and human sign-off. |
| Public Cloud GPU Inference | High egress costs for heavy DICOM image volumes and complex HIPAA data residency compliance barriers. | If on-premise hardware capital expenditure exceeds 3-year cloud reservations and BAA guarantees zero cross-tenant retention. |
| Monolithic Vision-Language Model | Single giant multi-modal model exceeds the 2,500 ms inference budget and increases hallucination risks. | Specialized multi-modal edge models achieving sub-1,000 ms latency with verified medical diagnostic certifications. |
## Contracts and Invariants
Clinical Model Weight Promotion Gate [INV-AIS-01]
Production model weights must never be updated autonomously. Any candidate weight release
must execute against the 5,000-study shadow evaluation suite, maintain F1 >= 0.94 without
negative drift on rare pulmonary classes, and receive cryptographic approval from Dr. Aris Thorne.
Low-Confidence Human Fallback [INV-AIS-02]
Whenever the model prediction confidence score for any detected finding is < 0.85, the platform
must abstain from autonomous diagnostic suggestions and immediately enqueue the DICOM study
to the senior radiologist urgent review queue with diagnostic flag `ALERT_MODEL_ABSTAIN`.
End-to-End Cryptographic Model Lineage [INV-AIS-03]
Every production inference result must record immutable metadata: `input_dicom_sha256`,
`model_weight_sha256`, `runtime_container_digest`, and `confidence_scores` persisted to
tamper-evident audit storage retained for 7 years.
Inference Latency SLA Ceiling [INV-AIS-04]
The GPU inference pipeline must return diagnostic output in p95 <= 2,500 ms under sustained
load of 15,000 studies/day (~10 studies/sec peak).
## Ownership and Handoffs
| Concern | Owner | Handoff payload | Blocked until |
|---|---|---|---|
| Clinical Validation & Weight Sign-Off | Chief Medical Officer (Dr. Aris Thorne) | `clinical_evaluation_report` | 5,000-study shadow benchmark passes |
| GPU Infrastructure & Triton Serving | AI Systems Engineering (Marcus Vance) | `gpu_cluster_provisioning_spec` | H100 node hardware acceptance |
| Data Governance & HIPAA Anonymization | Data Protection Officer | `dicom_deidentification_contract` | HIPAA compliance sign-off |
| PACS / DICOM Gateway Integration | Hospital Systems Integration Team | HL7/DICOM C-STORE receiver spec | Network firewall traversal |
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 15,000 DICOM studies daily volume | provided | Workload intake | Current |
| 4x NVIDIA H100 on-premise cluster | provided | Infrastructure intake | Current |
| Inference latency p95 <= 2,500 ms | provided | SLA constraint | Current |
| Rejection of autonomous retraining | decided | Dr. Aris Thorne (CMO) | 2026-09-15 |
| 5,000-study shadow gate with F1 >= 0.94 | provided | Clinical governance rule | Current |
| 0.85 confidence threshold fallback | decided | Architectural invariant INV-AIS-02 | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against AI system architecture contracts:
- **Governance Check**: PASS. Hard programmatic gate prohibits autonomous weight promotion; requires CMO approval.
- **Safety Fallback**: PASS. Abstention logic routes predictions < 0.85 to human radiologist queue.
- **Hardware Sizing**: PASS. 4x H100 nodes sized for 10 studies/sec burst throughput within 2,500 ms latency SLA.
- **Auditability**: PASS. 7-year SHA-256 lineage tracking for input DICOM, weights, and predictions specified.
## Open Decisions
- `DEC-AIS-01`: Dr. Thorne to verify whether pediatric chest studies require an independent distinct model classification head (Owner: Dr. Aris Thorne).
## Next steps
1. Marcus Vance deploys Triton Inference Server on the 4x NVIDIA H100 cluster and validates GPU memory allocation.
2. Data engineering team sets up the DICOM de-identification gateway ensuring zero PHI enters training datasets.
3. Establish the automated 5,000-case shadow evaluation runner in the private cloud staging environment.
ai-system-architect.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the integration architecture of an AI-enabled application across product behavior, model interaction, prompts/context, retrieval or memory, tools or agents, deterministic application components, safety and policy controls, evaluation, serving, operations and change lifecycle. It defines component responsibilities, end-to-end flows, cross-concern contracts, evidence boundaries, degraded behavior and release dependencies without taking over each specialist concern.
Use it when
- User or system inputs pass through deterministic validation, AI inference and application post-processing
- Model, prompt, context, retrieval, memory, tools or agents require explicit interaction contracts
- AI output can affect users, records, decisions or external systems and needs bounded authority
- Quality, latency, availability, privacy, safety and cost pressures must be traced to owner-supplied scenarios rather than generic trade-offs
- Fallback, abstention, human handoff or deterministic degradation spans multiple AI components
- Model/prompt/index/tool/policy revisions must remain compatible and observable as one release unit or coordinated set
For example: “We want AI in our expense tool: read a receipt photo, fill the claim, flag policy violations, and answer questions about the policy.”
What you get
- architecture/ai-system-architect/README.md
- architecture/ai-system-architect/00-overview/ai-system-architect-overview.md
- architecture/ai-system-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/provider-contract.md, {module}/translation.md, {module}/failure-mapping.md, {module}/credentials.md, {module}/idempotency.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use for model serving, prompt, RAG, memory, tool calling, agent, vector search, or evaluation design alone; traditional deterministic application architecture; implementing a named AI framework or provider; or requests triggered only by words such as AI, LLM, GenAI, chatbot, RAG, agent, model, copilot, multimodal, intelligent, quality, latency, or cost.
How it works
- Establish the deterministic baseline.
- Name the claims and who is harmed if they fail.
- Assign the minimum AI responsibility.
- Draw one canonical request flow.
- Fix the component contracts.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install