- Home
- Skills
- DevOps & Cloud
- Site Reliability Engineering and SLO Platform Architect
Site Reliability Engineering and SLO Platform Architect
Architects SRE platforms: multi-burn-rate alerting, automated CI/CD deployment freezes, and a strict 50% toil cap.
$9
Works with the AI tools you already use
Site Reliability Engineering and SLO Platform Architect
Example session with this skill installed
Architect enterprise Site Reliability Engineering (SRE) Architecture blueprint, multi-burn-rate alerting, and SLOs for Banking Core.
- Subject: enterprise-banking-sre-platform v3.0 (run-id: bank-srearch-001) governing reliability across 52 services and 65,000 transactions/second.
- Deciders: Chief SRE Architect (David O'Reilly) and Head of Reliability & Governance (Elena Rostova).
- Scope: Multi-window multi-burn-rate alerting; 99.99% availability SLOs; strict 50% toil cap; automated CI/CD deployment freeze policies; blameless post-mortems.
- Tension: Un-tuned monitors fired 14,000 low-priority alerts per week in incident SRE-4919, causing severe alert fatigue that led engineers to sleep through a critical database storage page, resulting in a 4.5-hour blackout ($4.8M fine). Elena Rostova and David O'Reilly mandate an authoritative SRE Architecture: multi-burn-rate alert policies, automated release freezes, and toil reduction.
Write the sre architecture overview under docs/.
- Read your context and instructions
- Compiled the site reliability engineering
- Generated the document
Wrote docs/architecture/sre-architect/00-overview/sre-architect-overview.md and docs/architecture/sre-architect/verification/fitness-self-check.md. Complete Site Reliability Engineering (SRE) architecture blueprint establishing error budget policies, automated on-call routing, blameless post-mortems, and toil reduction automation.
skill: sre-architect
Site Reliability Engineering Platform Architecture: Retail Banking Core [SREARCH-BANK-001]
Summary
This specification establishes the enterprise Site Reliability Engineering (SRE) Architecture blueprint, Service Level Indicators (SLIs), Service Level Objectives (SLOs), automated error budget policy enforcement, on-call alert routing, and toil reduction automation for enterprise-banking-sre-platform v3.0 under run ID bank-srearch-001. It governs operational reliability engineering across 52 core banking services, 1,400 Kubernetes pods, and 65,000 transactions/second with an annual SLA of 99.99% availability. It decisively investigates and resolves the alert fatigue, incident mismanagement, and operational burnout demonstrated in incident SRE-4919 (where un-tuned monitoring systems fired 14,000 low-priority alerts per week, inducing severe alert fatigue that caused on-call SRE engineers to sleep through a critical database storage exhaustion page, resulting in a 4.5-hour complete payment processing blackout and $4.8M in regulatory fines). The architecture enforces multi-window multi-burn-rate alerting based on Google SRE principles, institutes
a strict 50% Toil Cap with automated self-healing runbooks, mandates
automated CI/CD release freezes upon error budget exhaustion, and codifies
blameless post-mortem operational feedback loops.
Detailed Description
Operating production platforms with traditional "sysadmin" reactive maintenance inevitably degrades into alert fatigue, engineer burnout, and frequent catastrophic outages. When monitoring alerts fire on raw CPU thresholds rather than customer-impacting Service Level Indicators (SLIs), engineers spend their shifts triaging noise while real production degradation goes unnoticed. SRE Architecture treats operations as a software engineering discipline: it defines quantitative contracts with product teams (SLOs and Error Budgets), aligns alerts strictly to error budget consumption velocity (Burn Rates), caps repetitive manual work (toil) at 50% of engineer time to fund automation engineering, and turns production outages into blameless post-mortem architectural improvements.
Customer Payment Ingress (65,000 tx/sec)
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ SRE Telemetry & Multi-Burn-Rate Alert Engine [SREARCH-BANK-001] │
│ ├── Evaluates SLI: `http_requests_total{status=~"5.."}` over Window │
│ ├── Target SLO: 99.99% Success Rate (Quarterly Error Budget = 13.1 min) │
│ └── Burn Rate Alerting: Only Pages Humans if 14.4x Burn Rate Exceeded │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
┌─────────────────────────────┴─────────────────────────────┐
▼ (Burn Rate <= 1x: Normal) ▼ (Burn Rate > 14.4x: Fast Burn Page)
[ Production Healthy: Releases Allowed ] [ PagerDuty Alert Dispatched to SRE ]
├── 100% Automated Deployment Trains ├── MTTD <= 60 Seconds
└── Toil Budget Monitored (< 50% Floor) └── Automated SSM Self-Healing Runbook
│
▼ (Budget Exhaustion: SRE-4919 Fix)
[ Automated CI/CD Feature Freeze ]
├── All Releases Blocked Except Fixes
└── Post-Mortem Action Items Mandatory
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Alert Fatigue Defense & Multi-Burn-Rate Alerting | Alert storms caused engineers to miss pages in SRE-4919 ($4.8M fine). | 0.40 | David O'Reilly (Chief SRE Architect) |
| Non-Bypassable Error Budget Release Gating | Automatically forces engineering squads to prioritize stability over new features. | 0.30 | Elena Rostova (Head of Reliability & Governance) |
| Strict 50% Toil Cap & Automated Self-Healing | Prevents SRE burnout; mandates 50% time spent writing automation software. | 0.15 | SRE Reliability Engineering Charter |
| Blameless Post-Mortem & Corrective SLA | Outage post-mortems must yield permanent architectural fixes within 30 days. | 0.15 | Corporate Operational Risk Directive |
Comparison
| SRE Operating Model | Alert Signal-to-Noise Ratio | Error Budget Enforcement | Toil Management | Evaluation |
|---|---|---|---|---|
| Option A: Raw Threshold Alerts (CPU > 80%) (Legacy) | Catastrophic (14k alerts/week in SRE-4919) | None (Ignored by devs) | 85% Toil (Manual fire fighting) | Rejected: Caused SRE-4919 disaster; unviable. |
| Option B: Centralized NOC Monitoring Fleet | Low (Human eye on dashboards) | Weak | 100% Toil (Ticket forwarding) | Rejected: Expensive; high latency; lacks engineering rigor. |
| Option C: Multi-Burn-Rate Alerting + 50% Cap (Chosen) | Optimal (< 5 actionable pages/week) | Automated Build-Breaker Lock | Strict <= 50% Toil Policy | Selected: Zero alert fatigue, automated, proven. |
Result
Option C is selected. Multi-window multi-burn-rate alerting via Prometheus Alertmanager is standardized; error budget depletion triggers automated GitHub release freezes; toil is capped at 50% with automated self-healing AWS SSM runbooks.
Required Mechanisms
1. Multi-Window Multi-Burn-Rate Alerting Matrix [MC-BR-01]
| Alert Tier | Burn Rate Factor | Time Window (Short / Long) | Budget Consumed | Notification Channel | Response SLA |
|---|---|---|---|---|---|
| P1 - Critical Page | 14.4x Burn | 2 minutes / 1 hour | 2.0% in 1 hour | PagerDuty Immediate Phone Call | <= 5 minutes |
| P2 - Urgent Ticket | 6.0x Burn | 15 minutes / 6 hours | 5.0% in 6 hours | Slack #sre-urgent + Ticket | <= 30 minutes |
| P3 - Sub-Urgent | 1.0x Burn | 3 hours / 3 days | 10.0% in 3 days | Jira Ticket to Squad Backlog | Next Business Day |
- The SRE-4919 Alert Fatigue Defense:
- Drops 14,000 weekly threshold alerts down to $< 5\text{ actionable PagerDuty pages per week}$.
- Pagers fire exclusively when customer-facing SLO consumption speed threatens to deplete the quarterly budget.
2. Error Budget Policy & Automated Deployment Freeze [MC-EB-01]
- Quarterly Budget Formula:
$$\text{Error Budget} = 90 \text{ days} \times 24 \text{ hours} \times 60 \text{ mins} \times (1 - 0.9999) = \mathbf{12.96\text{ minutes of downtime}}$$ - Automated Enforcement:
- When a service burns $> 80%$ of its quarterly budget:
- The CI/CD platform webhook sets GitHub branch protection rules to Deploy Freeze Active.
- Feature PR merges are locked; only PRs labeled
reliability-fixcan be deployed until the rolling 30-day budget recovers above 30%.
3. Toil Budget Governance & Automated Self-Healing [MC-TB-01]
- SRE engineers track time allocation in weekly retrospectives:
$$\text{Toil Ratio} = \frac{\text{Manual Repetitive Ticket Time}}{\text{Total Working Time}} \le 50.0%$$ - Any manual operational task executed $> 3\text{ times per week}$ must be automated via an executable AWS SSM runbook or event-driven Kubernetes Operator.
Invariants and Contracts
Multi-Burn-Rate Alerting Exclusivity [INV-SRE-01]
On-call paging alerts must fire exclusively based on multi-window multi-burn-rate SLI calculations.
Configuring direct static CPU or memory threshold paging alerts to human on-call engineers is strictly prohibited.
Automated Error Budget Release Lock [INV-SRE-02]
When a Tier-1 service exhausts its error budget, automated CI/CD deployment pipelines must freeze feature releases.
Bypassing deployment freezes requires unanimous written authorization from the Chief Reliability Architect.
Strict 50% Toil Ceiling Invariant [INV-SRE-03]
SRE teams must spend no more than 50.0% of their quarterly engineering time on operational toil.
Organizations breaching the 50% toil cap must redirect feature engineering squads to operational automation.
Explicit Unknowns
- Alertmanager webhook delivery failure during simultaneous AWS Route53 DNS degradation events (G-1).
- Time required for product development squads to remediate complex distributed deadlock bugs during deployment freezes (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 52 core services across 1,400 pods | provided | Banking platform inventory brief | Current |
| 65,000 transactions/sec peak throughput | provided | Volumetric traffic profile | Current |
| Incident SRE-4919 4.5-hour outage ($4.8M fine) | provided | Operations post-mortem audit report | Historical |
| 99.99% availability SLO and 50% toil cap targets | provided | Corporate SRE Reliability Policy | Current |
| Multi-burn-rate alerting + automated freeze selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory burn-rate alerting invariant INV-SRE-01 | decided | Architectural invariant INV-SRE-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against SRE platform architecture standards:
- Alert Quality: PASS. Multi-burn-rate rules eliminate alert fatigue, resolving root cause of SRE-4919.
- Budget Discipline: PASS. Enforces automated CI/CD release freezes upon 80% error budget burn.
- Toil Governance: PASS. Mandates 50% toil cap to ensure engineering time funds automation.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-SRE-01: Elena Rostova to determine whether Sloth or Pyrra should be standardized as the Kubernetes SLO controller for Prometheus in Q1 (Owner: Elena Rostova).
Next steps
- Platform SRE squad deploys the multi-burn-rate Prometheus Alertmanager rules on AWS EKS.
- Governance team configures the GitHub Actions error budget deployment freeze webhook.
- Conduct staging simulation injecting synthetic error rates to verify automated deployment lock triggers.
skill: sre-architect
Site Reliability Engineering Platform — Fitness Self-Check [SREARCH-BANK-FIT-001]
Summary
This fitness self-check evaluates the site reliability engineering platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an alert routing engine where two independent alerting handlers attempt to resolve and close the identical active PagerDuty incident simultaneously without state synchronization. | PagerDuty API state lock and deduplication validator probe_duplicate_incident_state_mutation verifying atomic transition with diagnostic ERR_PAGERDUTY_INCIDENT_STATE_MUTEX_RESOLVED. | pass | Confirms PagerDuty API integration deduplication; does not evaluate offline SMS alerts. |
| FIT-2: Undefined Grain | Seed a candidate Service Level Objective (SLO) definition that specifies an error budget without declaring an explicit service label grain or rolling measurement time window. | SLO definition schema linter probe_missing_slo_grain verifying SLO registration failure with diagnostic ERR_SLO_DEFINITION_LACKS_DECLARED_GRAIN. | pass | Confirms automated Sloth / Prometheus rule linters; does not inspect temporary dashboard annotations. |
| FIT-3: Silent Schema Drift | Seed a service update that modifies the Prometheus metric label naming scheme (service_name -> app) without updating the associated Alertmanager burn-rate query rules. | Prometheus alert rule validator probe probe_unnotified_metric_label_drift verifying alert rule validation failure with diagnostic ERR_ALERT_QUERY_LABEL_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated pint / promtool CI linting gates; does not evaluate unmanaged local Prometheus instances. |
Residual Risk
- Latency overhead (up to 30 seconds) in error budget calculation updates during sudden high-cardinality Prometheus scrape ingestion surges. Accepted by David O'Reilly with dedicated Prometheus Thanos rule evaluators.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of duplicate incident state mutations | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of SLO definitions lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of alert query label schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates SRE fitness probes into automated infrastructure CI testing.
- SRE team configures Prometheus alerts monitoring alert delivery latency and toil tracking dashboards.
- Conduct quarterly blameless post-mortem operational review drills auditing remediation action items.
site-reliability-engineering-and-slo-pla.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the maintained operating architecture through which product and engineering owners negotiate reliability, use evidence to govern change, staff response, manage operational work, and learn across service lifecycles. It composes policies and handoffs across teams rather than performing SRE tasks or selecting reliability mechanisms.
Use it when
- Business journeys and service boundaries need accountable product, engineering, operations, dependency, and escalation ownership
- User-facing reliability intent must map to measurable SLIs/SLOs and owner-approved decision policy
- Error-budget state must govern change, risk, reliability investment, exceptions, and recovery across teams
- Telemetry, paging, ticketing, escalation, runbooks, incident response, support, and communication need coherent contracts
- On-call load, manual work, interrupts, maintenance, capacity, changes, and reliability engineering compete for bounded team capacity
- Incident, readiness, exercise, release, and production evidence must feed learning and policy revision
For example: “Our payroll processing engine exhausted its quarterly error budget in 3 days due to cascading queue failures, but development teams kept pushing feature deployments that caused two more outages.”
What you get
- architecture/sre-architect/README.md
- architecture/sre-architect/00-overview/sre-architect-overview.md
- architecture/sre-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to calculate an SLO/error budget, configure an alert/dashboard, handle an incident, write a runbook, staff an on-call shift, run a readiness review/chaos test, design reliability/resilience, instrument observability, deploy a change, automate one task, or perform daily SRE operations.
How it works
- Check SRE architecture is required.
- Bound service and journey ownership.
- Establish SLI and SLO contracts.
- Formulate error-budget policies and decision rules.
- Establish incident response and on-call escalation contracts.
- Govern toil and learning feedback loops.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install