Site Reliability Engineering and SLO Platform Architect

    1

    Architects SRE platforms: multi-burn-rate alerting, automated CI/CD deployment freezes, and a strict 50% toil cap.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Site Reliability Engineering and SLO Platform Architect

    Example session with this skill installed

    Architect enterprise Site Reliability Engineering (SRE) Architecture blueprint, multi-burn-rate alerting, and SLOs for Banking Core.

    • Subject: enterprise-banking-sre-platform v3.0 (run-id: bank-srearch-001) governing reliability across 52 services and 65,000 transactions/second.
    • Deciders: Chief SRE Architect (David O'Reilly) and Head of Reliability & Governance (Elena Rostova).
    • Scope: Multi-window multi-burn-rate alerting; 99.99% availability SLOs; strict 50% toil cap; automated CI/CD deployment freeze policies; blameless post-mortems.
    • Tension: Un-tuned monitors fired 14,000 low-priority alerts per week in incident SRE-4919, causing severe alert fatigue that led engineers to sleep through a critical database storage page, resulting in a 4.5-hour blackout ($4.8M fine). Elena Rostova and David O'Reilly mandate an authoritative SRE Architecture: multi-burn-rate alert policies, automated release freezes, and toil reduction.

    Write the sre architecture overview under docs/.

    • Read your context and instructions
    • Compiled the site reliability engineering
    • Generated the document

    Wrote docs/architecture/sre-architect/00-overview/sre-architect-overview.md and docs/architecture/sre-architect/verification/fitness-self-check.md. Complete Site Reliability Engineering (SRE) architecture blueprint establishing error budget policies, automated on-call routing, blameless post-mortems, and toil reduction automation.


    skill: sre-architect

    Site Reliability Engineering Platform Architecture: Retail Banking Core [SREARCH-BANK-001]

    Summary

    This specification establishes the enterprise Site Reliability Engineering (SRE) Architecture blueprint, Service Level Indicators (SLIs), Service Level Objectives (SLOs), automated error budget policy enforcement, on-call alert routing, and toil reduction automation for enterprise-banking-sre-platform v3.0 under run ID bank-srearch-001. It governs operational reliability engineering across 52 core banking services, 1,400 Kubernetes pods, and 65,000 transactions/second with an annual SLA of 99.99% availability. It decisively investigates and resolves the alert fatigue, incident mismanagement, and operational burnout demonstrated in incident SRE-4919 (where un-tuned monitoring systems fired 14,000 low-priority alerts per week, inducing severe alert fatigue that caused on-call SRE engineers to sleep through a critical database storage exhaustion page, resulting in a 4.5-hour complete payment processing blackout and $4.8M in regulatory fines). The architecture enforces multi-window multi-burn-rate alerting based on Google SRE principles, institutes

    a strict 50% Toil Cap with automated self-healing runbooks, mandates

    automated CI/CD release freezes upon error budget exhaustion, and codifies

    blameless post-mortem operational feedback loops.

    Detailed Description

    Operating production platforms with traditional "sysadmin" reactive maintenance inevitably degrades into alert fatigue, engineer burnout, and frequent catastrophic outages. When monitoring alerts fire on raw CPU thresholds rather than customer-impacting Service Level Indicators (SLIs), engineers spend their shifts triaging noise while real production degradation goes unnoticed. SRE Architecture treats operations as a software engineering discipline: it defines quantitative contracts with product teams (SLOs and Error Budgets), aligns alerts strictly to error budget consumption velocity (Burn Rates), caps repetitive manual work (toil) at 50% of engineer time to fund automation engineering, and turns production outages into blameless post-mortem architectural improvements.

    Customer Payment Ingress (65,000 tx/sec)
                             │
                             ▼
    ┌─────────────────────────────────────────────────────────────────────────────┐
    │ SRE Telemetry & Multi-Burn-Rate Alert Engine [SREARCH-BANK-001]             │
    │   ├── Evaluates SLI: `http_requests_total{status=~"5.."}` over Window       │
    │   ├── Target SLO: 99.99% Success Rate (Quarterly Error Budget = 13.1 min)   │
    │   └── Burn Rate Alerting: Only Pages Humans if 14.4x Burn Rate Exceeded     │
    └──────────────────────────────────────┬──────────────────────────────────────┘
                                           │
             ┌─────────────────────────────┴─────────────────────────────┐
             ▼ (Burn Rate <= 1x: Normal)                                 ▼ (Burn Rate > 14.4x: Fast Burn Page)
    [ Production Healthy: Releases Allowed ]                    [ PagerDuty Alert Dispatched to SRE ]
      ├── 100% Automated Deployment Trains                       ├── MTTD <= 60 Seconds
      └── Toil Budget Monitored (< 50% Floor)                     └── Automated SSM Self-Healing Runbook
                                                                         │
                                                                         ▼ (Budget Exhaustion: SRE-4919 Fix)
                                                                [ Automated CI/CD Feature Freeze ]
                                                                  ├── All Releases Blocked Except Fixes
                                                                  └── Post-Mortem Action Items Mandatory
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Alert Fatigue Defense & Multi-Burn-Rate AlertingAlert storms caused engineers to miss pages in SRE-4919 ($4.8M fine).0.40David O'Reilly (Chief SRE Architect)
    Non-Bypassable Error Budget Release GatingAutomatically forces engineering squads to prioritize stability over new features.0.30Elena Rostova (Head of Reliability & Governance)
    Strict 50% Toil Cap & Automated Self-HealingPrevents SRE burnout; mandates 50% time spent writing automation software.0.15SRE Reliability Engineering Charter
    Blameless Post-Mortem & Corrective SLAOutage post-mortems must yield permanent architectural fixes within 30 days.0.15Corporate Operational Risk Directive

    Comparison

    SRE Operating ModelAlert Signal-to-Noise RatioError Budget EnforcementToil ManagementEvaluation
    Option A: Raw Threshold Alerts (CPU > 80%) (Legacy)Catastrophic (14k alerts/week in SRE-4919)None (Ignored by devs)85% Toil (Manual fire fighting)Rejected: Caused SRE-4919 disaster; unviable.
    Option B: Centralized NOC Monitoring FleetLow (Human eye on dashboards)Weak100% Toil (Ticket forwarding)Rejected: Expensive; high latency; lacks engineering rigor.
    Option C: Multi-Burn-Rate Alerting + 50% Cap (Chosen)Optimal (< 5 actionable pages/week)Automated Build-Breaker LockStrict <= 50% Toil PolicySelected: Zero alert fatigue, automated, proven.

    Result

    Option C is selected. Multi-window multi-burn-rate alerting via Prometheus Alertmanager is standardized; error budget depletion triggers automated GitHub release freezes; toil is capped at 50% with automated self-healing AWS SSM runbooks.


    Required Mechanisms

    1. Multi-Window Multi-Burn-Rate Alerting Matrix [MC-BR-01]
    Alert TierBurn Rate FactorTime Window (Short / Long)Budget ConsumedNotification ChannelResponse SLA
    P1 - Critical Page14.4x Burn2 minutes / 1 hour2.0% in 1 hourPagerDuty Immediate Phone Call<= 5 minutes
    P2 - Urgent Ticket6.0x Burn15 minutes / 6 hours5.0% in 6 hoursSlack #sre-urgent + Ticket<= 30 minutes
    P3 - Sub-Urgent1.0x Burn3 hours / 3 days10.0% in 3 daysJira Ticket to Squad BacklogNext Business Day
    • The SRE-4919 Alert Fatigue Defense:
      • Drops 14,000 weekly threshold alerts down to $< 5\text{ actionable PagerDuty pages per week}$.
      • Pagers fire exclusively when customer-facing SLO consumption speed threatens to deplete the quarterly budget.
    2. Error Budget Policy & Automated Deployment Freeze [MC-EB-01]
    • Quarterly Budget Formula:
      $$\text{Error Budget} = 90 \text{ days} \times 24 \text{ hours} \times 60 \text{ mins} \times (1 - 0.9999) = \mathbf{12.96\text{ minutes of downtime}}$$
    • Automated Enforcement:
      • When a service burns $> 80%$ of its quarterly budget:
      • The CI/CD platform webhook sets GitHub branch protection rules to Deploy Freeze Active.
      • Feature PR merges are locked; only PRs labeled reliability-fix can be deployed until the rolling 30-day budget recovers above 30%.
    3. Toil Budget Governance & Automated Self-Healing [MC-TB-01]
    • SRE engineers track time allocation in weekly retrospectives:
      $$\text{Toil Ratio} = \frac{\text{Manual Repetitive Ticket Time}}{\text{Total Working Time}} \le 50.0%$$
    • Any manual operational task executed $> 3\text{ times per week}$ must be automated via an executable AWS SSM runbook or event-driven Kubernetes Operator.

    Invariants and Contracts

    Multi-Burn-Rate Alerting Exclusivity [INV-SRE-01]
      On-call paging alerts must fire exclusively based on multi-window multi-burn-rate SLI calculations.
      Configuring direct static CPU or memory threshold paging alerts to human on-call engineers is strictly prohibited.
    
    Automated Error Budget Release Lock [INV-SRE-02]
      When a Tier-1 service exhausts its error budget, automated CI/CD deployment pipelines must freeze feature releases.
      Bypassing deployment freezes requires unanimous written authorization from the Chief Reliability Architect.
    
    Strict 50% Toil Ceiling Invariant [INV-SRE-03]
      SRE teams must spend no more than 50.0% of their quarterly engineering time on operational toil.
      Organizations breaching the 50% toil cap must redirect feature engineering squads to operational automation.
    

    Explicit Unknowns

    • Alertmanager webhook delivery failure during simultaneous AWS Route53 DNS degradation events (G-1).
    • Time required for product development squads to remediate complex distributed deadlock bugs during deployment freezes (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    52 core services across 1,400 podsprovidedBanking platform inventory briefCurrent
    65,000 transactions/sec peak throughputprovidedVolumetric traffic profileCurrent
    Incident SRE-4919 4.5-hour outage ($4.8M fine)providedOperations post-mortem audit reportHistorical
    99.99% availability SLO and 50% toil cap targetsprovidedCorporate SRE Reliability PolicyCurrent
    Multi-burn-rate alerting + automated freeze selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory burn-rate alerting invariant INV-SRE-01decidedArchitectural invariant INV-SRE-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against SRE platform architecture standards:

    • Alert Quality: PASS. Multi-burn-rate rules eliminate alert fatigue, resolving root cause of SRE-4919.
    • Budget Discipline: PASS. Enforces automated CI/CD release freezes upon 80% error budget burn.
    • Toil Governance: PASS. Mandates 50% toil cap to ensure engineering time funds automation.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-SRE-01: Elena Rostova to determine whether Sloth or Pyrra should be standardized as the Kubernetes SLO controller for Prometheus in Q1 (Owner: Elena Rostova).

    Next steps

    1. Platform SRE squad deploys the multi-burn-rate Prometheus Alertmanager rules on AWS EKS.
    2. Governance team configures the GitHub Actions error budget deployment freeze webhook.
    3. Conduct staging simulation injecting synthetic error rates to verify automated deployment lock triggers.

    skill: sre-architect

    Site Reliability Engineering Platform — Fitness Self-Check [SREARCH-BANK-FIT-001]

    Summary

    This fitness self-check evaluates the site reliability engineering platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed an alert routing engine where two independent alerting handlers attempt to resolve and close the identical active PagerDuty incident simultaneously without state synchronization.PagerDuty API state lock and deduplication validator probe_duplicate_incident_state_mutation verifying atomic transition with diagnostic ERR_PAGERDUTY_INCIDENT_STATE_MUTEX_RESOLVED.passConfirms PagerDuty API integration deduplication; does not evaluate offline SMS alerts.
    FIT-2: Undefined GrainSeed a candidate Service Level Objective (SLO) definition that specifies an error budget without declaring an explicit service label grain or rolling measurement time window.SLO definition schema linter probe_missing_slo_grain verifying SLO registration failure with diagnostic ERR_SLO_DEFINITION_LACKS_DECLARED_GRAIN.passConfirms automated Sloth / Prometheus rule linters; does not inspect temporary dashboard annotations.
    FIT-3: Silent Schema DriftSeed a service update that modifies the Prometheus metric label naming scheme (service_name -> app) without updating the associated Alertmanager burn-rate query rules.Prometheus alert rule validator probe probe_unnotified_metric_label_drift verifying alert rule validation failure with diagnostic ERR_ALERT_QUERY_LABEL_SCHEMA_DRIFT_DETECTED.passConfirms automated pint / promtool CI linting gates; does not evaluate unmanaged local Prometheus instances.

    Residual Risk

    • Latency overhead (up to 30 seconds) in error budget calculation updates during sudden high-cardinality Prometheus scrape ingestion surges. Accepted by David O'Reilly with dedicated Prometheus Thanos rule evaluators.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of duplicate incident state mutationsderivedFIT-1 probe result2026-09-15
    Rejection of SLO definitions lacking declared grainderivedFIT-2 probe result2026-09-15
    Rejection of alert query label schema driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates SRE fitness probes into automated infrastructure CI testing.
    2. SRE team configures Prometheus alerts monitoring alert delivery latency and toil tracking dashboards.
    3. Conduct quarterly blameless post-mortem operational review drills auditing remediation action items.

    site-reliability-engineering-and-slo-pla.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Map business journeys to service ownership and escalation pathsDefine queryable SLI contracts and SLO targets for user-facing servicesEstablish error-budget policies and automated deployment freeze rulesGovern team toil limits and incident learning feedback loops

    About this skill

    What it does

    This skill owns the maintained operating architecture through which product and engineering owners negotiate reliability, use evidence to govern change, staff response, manage operational work, and learn across service lifecycles. It composes policies and handoffs across teams rather than performing SRE tasks or selecting reliability mechanisms.

    Use it when

    • Business journeys and service boundaries need accountable product, engineering, operations, dependency, and escalation ownership
    • User-facing reliability intent must map to measurable SLIs/SLOs and owner-approved decision policy
    • Error-budget state must govern change, risk, reliability investment, exceptions, and recovery across teams
    • Telemetry, paging, ticketing, escalation, runbooks, incident response, support, and communication need coherent contracts
    • On-call load, manual work, interrupts, maintenance, capacity, changes, and reliability engineering compete for bounded team capacity
    • Incident, readiness, exercise, release, and production evidence must feed learning and policy revision

    For example: “Our payroll processing engine exhausted its quarterly error budget in 3 days due to cascading queue failures, but development teams kept pushing feature deployments that caused two more outages.”

    What you get

    • architecture/sre-architect/README.md
    • architecture/sre-architect/00-overview/sre-architect-overview.md
    • architecture/sre-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to calculate an SLO/error budget, configure an alert/dashboard, handle an incident, write a runbook, staff an on-call shift, run a readiness review/chaos test, design reliability/resilience, instrument observability, deploy a change, automate one task, or perform daily SRE operations.

    How it works

    1. Check SRE architecture is required.
    2. Bound service and journey ownership.
    3. Establish SLI and SLO contracts.
    4. Formulate error-budget policies and decision rules.
    5. Establish incident response and on-call escalation contracts.
    6. Govern toil and learning feedback loops.
    7. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions