- Home
- Skills
- DevOps & Cloud
- SLI, SLO and Error Budget Design
SLI, SLO and Error Budget Design
Defines SLIs, SLOs and error budgets from real user journeys: what counts as good, targets, windows, budget math.
$5
Works with the AI tools you already use
SLI, SLO and Error Budget Design
Example session with this skill installed
Design the SLI, SLO, and Error Budget contract for our Retail Checkout journey.
- Subject: checkout-service v3.4.0, critical user journey (CUJ): Customer Places Order via POST /v1/checkout/orders (run-id: checkout-slo-001).
- Deciders: VP of Reliability Engineering (David O'Reilly) and VP of Product (Rachel Adams).
- Traffic Profile: 45 million checkout attempts per 28-day rolling window (~18.6 orders/sec average, peak 120 orders/sec).
- Baseline & Target: Historical availability is 99.78%; Product wants 99.99% SLO, but SRE warns that our third-party payment gateway SLA is only 99.95%, making 99.99% impossible without decoupling. Agreed compromise target: 99.9% availability and 99.0% latency <= 1,200 ms.
- Tension: Product insists on counting HTTP 422 validation errors as SLO failures; SRE insists client errors (4xx except 429) must be excluded from counts so bad client input does not exhaust the reliability budget.
- Error Budget Policy: 28-day rolling window; multi-window burn rate alerts (14.4x over 1h page on-call, 6x over 6h ticket); freeze feature releases if 100% budget burned until reliability actions merge.
Write the SLI/SLO contract specification under docs/.
- Read your context and instructions
- Compiled the sli slo
- Generated the UI component
Wrote docs/architecture/tasks/checkout-slo-001/sli-slo-design/sli-slo-contract.md. Complete SLI/SLO contract establishing availability and latency indicators for the retail checkout user journey, multi-window burn rate alerts, and error budget governance.
---
skill: sli-slo-design
---
# SLI, SLO and Error Budget Contract: Retail Checkout Journey [SLO-CHK-001]
## Summary
This contract establishes the Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budget governance policies for the Retail Checkout Critical User Journey (`POST /v1/checkout/orders`) on `checkout-service v3.4.0` under run ID `checkout-slo-001`. Evaluating 45 million order attempts across a 28-day rolling window, it resolves the dispute between Product and SRE by adopting an achievable 99.9% availability target (bounded by third-party payment gateway dependencies at 99.95%) and a 99.0% latency target (<= 1,200 ms). It filters out client validation errors (HTTP 4xx except 429) from indicator denominators, specifies multi-window multi-burn-rate alerting, and enacts a binding feature freeze policy upon budget exhaustion.
## Detailed Description
Retail checkout is the primary revenue-generating flow. Unrealistic SLO targets (such as 99.99%) induce alert fatigue and trigger unjustified feature freezes when upstream third-party payment rails degrade within their own contract. Conversely, failing to filter invalid client requests penalizes development teams for user input errors.
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| User-Centric Accuracy | SLIs must measure actual customer completion, not internal host CPU or synthetic pings. | 0.35 | David O'Reilly (VP SRE) |
| Dependency Feasibility | SLO targets cannot exceed the mathematical availability ceiling of critical upstream third-party payment gateways (99.95%). | 0.30 | Intake dependency constraint |
| Actionable Alert Signal-to-Noise | Burn rate alerts must page on-call engineers only when budget depletion threatens the 28-day target. | 0.20 | Google SRE alerting standard |
| Enforceable Product Governance | When the error budget is exhausted, engineering priorities must legally pivot to reliability remediation. | 0.15 | Rachel Adams (VP Product) |
### Comparison
| Dimension | Proposal A (Product Initial) | Proposal B (Chosen Standard) | Architectural Justification |
|---|---|---|---|
| Availability SLO Target | 99.99% (52 min downtime / 28 days) | 99.90% (40.3 min downtime / 28 days) | Upstream payment gateway SLA is 99.95%; composite availability cannot exceed 99.95% without async escrow. |
| Denominator Scope | All incoming HTTP requests including 4xx | Valid requests only (excludes 4xx except 429) | Bad user input (e.g. expired cards, malformed JSON) is not a service reliability failure. |
| Latency Boundary | p95 <= 500 ms | 99.0% <= 1,200 ms | Tail latency accommodates synchronous payment gateway authorization handshakes. |
### Result
Proposal B is selected. The service commits to two core SLOs across rolling 28-day windows:
1. **Availability SLO**: >= 99.90%
2. **Latency SLO**: >= 99.00% responses <= 1,200 ms
---
### Required Mechanisms
#### 1. Critical User Journey (CUJ) [MC-CUJ-01]
- **Journey**: Customer submits cart items and payment details to place an order.
- **Boundary**: Ingress API Gateway handling `POST /v1/checkout/orders`.
- **Classification**:
- *Valid Requests*: HTTP 2xx, 3xx, 5xx, and 429.
- *Excluded Requests*: HTTP 400, 401, 403, 404, 422 (client authentication and validation failures).
#### 2. Service Level Indicators (SLI) Formulation [MC-SLI-01]
##### SLI-1: Availability
$$\text{SLI}_{\text{avail}} = \frac{\sum \text{Requests with status } \in [200, 201, 202, 204]}{\sum \text{Valid Requests (2xx, 3xx, 5xx, 429)}}$$
##### SLI-2: Latency
$$\text{SLI}_{\text{latency}} = \frac{\sum \text{Successful Checkout Requests with duration } \le 1,200\text{ ms}}{\sum \text{Successful Checkout Requests (2xx)}}$$
#### 3. Error Budget Mathematics [MC-EB-01]
- **Rolling Window**: 28 days ($28 \times 24 \times 3600 = 2,419,200$ seconds).
- **Total Valid Events**: 45,000,000 checkout requests per window.
- **Availability Error Budget**:
$$\text{Budget} = 1 - 0.9990 = 0.0010 \quad (0.10\%)$$
$$\text{Allowed Failed Events} = 45,000,000 \times 0.0010 = 45,000\text{ failed orders / 28 days}$$
- **Latency Error Budget**:
$$\text{Budget} = 1 - 0.9900 = 0.0100 \quad (1.00\%)$$
$$\text{Allowed Slow Events} = 45,000,000 \times 0.0100 = 450,000\text{ slow orders / 28 days}$$
#### 4. Multi-Window Multi-Burn-Rate Alerting [MC-AL-01]
| Severity | Burn Rate Factor | Budget Consumed | Short Window | Long Window | Notification Channel |
|---|---|---|---|---|---|
| Page (P1) | 14.4x | 2.0% in 1 hour | 5 minutes | 1 hour | PagerDuty SRE On-Call |
| Page (P2) | 6.0x | 5.0% in 6 hours | 30 minutes | 6 hours | PagerDuty Primary Dev |
| Ticket (P3) | 1.0x | 10.0% in 3 days | 2 hours | 3 days | Jira Reliability Backlog |
*Alert condition requires BOTH short and long windows to exceed burn rate threshold to prevent false-positive alert flapping on transient spikes.*
#### 5. Error Budget Policy & Governance [MC-GOV-01]
- **Budget State Transitions**:
- **Healthy (Remaining > 20%)**: Standard feature deployments permitted without SRE review.
- **Warning (Remaining <= 20%)**: High-risk infrastructure changes deferred; daily standup budget reviews.
- **Exhausted (Remaining <= 0%)**: Binding feature freeze enacted:
1. Automated CI/CD deployment pipelines lock non-reliability pull requests.
2. All sprint capacity redirected to reliability remediation items.
3. Production deployments restricted to P0 security patches and SLO-restoring bug fixes.
4. Freeze lifts only when 28-day rolling budget recovers above 20%.
---
### Invariants and Contracts
Client Error Exclusion Invariant [INV-SLO-01]
Client validation errors (HTTP 400, 401, 403, 404, 422) must be excluded from both
the numerator and denominator of the availability SLI. Only server errors (5xx)
and rate-limiting shedding (429) count as negative reliability events.
Multi-Window Alerting Requirement [INV-SLO-02]
PagerDuty alerts must require concurrent threshold breaches across both a short window
and a long window. Single-point or instantaneous rate spikes are forbidden from paging.
Binding Feature Freeze Enforcement [INV-SLO-03]
Deployment pipeline release locks upon 100% error budget consumption are automated
and cannot be overridden without written joint sign-off from both Rachel Adams and David O'Reilly.
## Explicit Unknowns
- Third-party payment gateway maintenance downtime notification lead times (G-1).
- Telemetry sampling rate overhead on Prometheus / OpenTelemetry collectors at 120 orders/sec peak (G-2).
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 45M requests per 28-day window | provided | Traffic profile | Current |
| Third-party gateway SLA 99.95% | provided | Intake constraint | Current |
| 99.90% availability target | decided | Rachel Adams & David O'Reilly | 2026-09-15 |
| 99.00% latency target <= 1,200 ms | decided | Rachel Adams & David O'Reilly | 2026-09-15 |
| Client 4xx exclusion from denominator | decided | SRE standard invariant INV-SLO-01 | 2026-09-15 |
| 14.4x 1h and 6x 6h burn rate thresholds | derived | Google SRE Multi-Window Math | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against SLI/SLO domain contracts:
- **Feasibility Check**: PASS. 99.90% target is mathematically achievable beneath 99.95% payment gateway ceiling.
- **Formula Exactness**: PASS. 45M * 0.0010 = 45,000 allowed failures; SLI math explicitly excludes 4xx.
- **Alert Soundness**: PASS. Multi-window burn rates prevent paging on transient 1-minute bursts.
- **Governance Clarity**: PASS. Concrete freeze triggers and release-blocking rules defined.
## Open Decisions
- `DEC-SLO-01`: SRE team to configure Prometheus recording rules in `monitoring/prometheus/slo_rules.yml` (Owner: David O'Reilly).
## Next steps
1. Rachel Adams (Product) and David O'Reilly (SRE) sign off on the error budget policy.
2. Implement Prometheus recording rules for 5-minute, 1-hour, and 6-hour request rate windows.
3. Configure GitHub Actions deployment gate to query error budget status before production CD promotions.
sli-slo-and-error-budget-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted user/service outcomes and reliability intent into bounded indicators, objectives and budget semantics. It defines what is measured and judged without choosing instrumentation, alerts, dashboards or external remedies.
Use it when
Use when an owned user journey or service capability needs an internal measurable reliability target and review contract.
For example: “Our executive team wants 'four nines' across the board, but our prescription refill deployment froze for two weeks because an internal background report generator ran out of error budget.”
What you get
- SLO / SLA Contract Spec
- Error Budget Report Template
- Burn Rate Alert Rules
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/sli-slo-design/.
What it will not do
Do not use for external SLA/remedies, metric instrumentation, dashboard/alert/product configuration or generic KPI selection.
How it works
- Check SLI/SLO contract is required.
- Identify critical user journeys and service capabilities.
- Formulate exact SLI numerator and denominator definitions.
- Establish target SLO percentages and evaluation windows.
- Calculate error budget allocation and burn rate alert policies.
- Write the SLI/SLO contract spec under <output_root>/architecture/tasks/{run-id}/sli-slo-design/slo-sla-contract-spec.md.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 13 days ago
- Passed all security checks, Safe to install