- Home
- Skills
- DevOps & Cloud
- Actionable Alert Contract Design
Actionable Alert Contract Design
Designs actionable alerting contracts: symptom-based triggers, routing matrices, deduplication, and runbook bindings.
$5
Works with the AI tools you already use
Actionable Alert Contract Design
Example session with this skill installed
Design actionable alerting and routing contracts for our Payments Processing Engine.
- Subject: payment-core v3.2.0 (run-id: payments-alert-001) processing 1,400 peak transactions/sec on Kubernetes across 3 AWS AZs.
- Deciders: Lead SRE Architect (Marcus Vance) and Payment Reliability Lead (Elena Rostova).
- Incidents: In INC-3118, a transient network blip triggered 4,200 individual alerts across 14 Slack channels and paged 8 engineers, causing severe alert fatigue that delayed incident resolution by 35 minutes.
- Tension: Dev team proposed alerts for every internal CPU spike (> 80%), memory threshold (> 75%), and exception log; Marcus Vance strictly forbids cause-based CPU alerts lacking user impact, mandating symptom-based alerting (multi-window SLO error burn rates, authorization p95 latency breaches), Alertmanager alert grouping and inhibition (cluster-down silences pod alerts), a strict 3-tier severity matrix (P1 page on-call, P2 business-hours ticket, P3 informational log), and mandatory runbook URLs (https://runbooks.internal/payments/...) on all paging alerts.
Write the alerting policy spec under docs/.
- Read your context and instructions
- Compiled the actionable alert contract
- Generated the UI component
Wrote docs/architecture/tasks/payments-alert-001/alerting-design/alerting-policy-spec.md. Complete actionable alerting contract specifying symptom-based SLO triggers, alert grouping, inhibition hierarchies, on-call escalation matrices, and mandatory runbook bindings.
---
skill: alerting-design
---
# Alerting Policy Spec: Core Payments Processing [ALERT-PAY-001]
## Summary
This specification establishes the actionable alerting contract, severity taxonomy, inhibition rules, and routing matrix for `payment-core v3.2.0` under run ID `payments-alert-001`, monitoring 1,400 peak payment transactions/second. It resolves the severe alert fatigue and noise demonstrated in incident INC-3118 (where a transient network glitch generated 4,200 simultaneous Slack alerts and paged 8 engineers). The design strictly bans cause-based noise (such as raw CPU > 80% alerts), mandating symptom-based alerting focused on customer experience (multi-window SLO error budget burn rates, authorization p95 latency breaches), Alertmanager grouping with inhibition trees, a 3-tier severity hierarchy, and mandatory validated runbook links (`https://runbooks.internal/payments/...`) on every paging alert.
## Detailed Description
Uncalibrated cause-based alerting produces alert fatigue, noise blindness, and delayed incident recovery. Paging engineers on ephemeral host metrics that have zero impact on customer transactions leads to missed P1 outages. This contract enforces Google SRE symptom-based alerting: an alert must represent an urgent, actionable problem requiring human intervention, with a documented response path.
Ingress Telemetry Stream (1,400 TPS)
│
▼
[ Prometheus Evaluation Engine (15s Interval) ]
├── Check 1: Multi-Window SLO Burn Rate (14.4x 1h / 6x 6h)
└── Check 2: Transaction Latency p95 > 180 ms for 3m
│
▼ (Alert Condition Tripped)
[ Alertmanager Routing & Grouping Core ]
├── Grouping: group_by: [alertname, cluster, service]
├── Inhibition: If ClusterNetworkDown is firing ──► Inhibit PodCrashLoop & InstanceDown
└── Routing:
├── Severity: P1 (Critical) ──► PagerDuty @payment-sre-oncall (Mandatory Runbook)
├── Severity: P2 (Warning) ──► Jira Service Desk (Response SLA: 4h)
└── Severity: P3 (Info) ──► Datadog Diagnostic Dashboard Log
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Alert Actionability & Noise Elimination | Paging engineers for transient, non-actionable blips induces alert fatigue and delayed incident response (INC-3118). | 0.40 | Marcus Vance (Lead SRE) |
| Customer-Impact Symptom Alignment | Alerts must track user-facing pain (failed payments, latency) rather than ephemeral node CPU spikes. | 0.30 | Elena Rostova (Payment Reliability) |
| Fast Incident Triage (MTTR Reduction) | Responders must be directed to exact operational runbooks within 30 seconds of receiving a page. | 0.15 | SRE Operations SLA |
| Alert Volume Bounding via Grouping | Aggregating related pod alerts into a single notification prevents Slack and pager floods. | 0.15 | Incident Response Policy |
### Comparison
| Alerting Strategy Candidate | Evaluation Target | Notification Volume under Blip | Responder Action Clarity | Evaluation |
|---|---|---|---|---|
| Option A: Raw Infrastructure Triggers | CPU > 80%, Disk > 75%, Exceptions | Catastrophic: 4,200 alerts across 14 channels | Poor: Unclear if customers are affected | Rejected: Triggered INC-3118 responder paralysis. |
| Option B: Single Static Error Rate (> 1%) | Instantaneous error percentage | High: False alarms on 10-second traffic dips | Moderate | Rejected: Lacks multi-window burn rate smoothing. |
| Option C: Multi-Window SLO Burn + Inhibition (Chosen) | 14.4x 1h / 6x 6h burn + Latency p95 | Single grouped notification per incident | High: Direct link to verified runbook | Selected: High signal-to-noise ratio, mathematically sound. |
### Result
Option C is selected. Prometheus evaluates multi-window SLO burn rates; Alertmanager groups and inhibits downstream noise, routing only actionable alerts to on-call engineers.
---
### Required Mechanisms
#### 1. Severity Classification Taxonomy [MC-ST-01]
| Severity Level | Definition & Customer Impact | Response Channel | Page SLA | Required Artifact |
|---|---|---|---|---|
| **P1 - Critical** | Severe transaction failure rate (> 1% drop) or complete service unavailability | PagerDuty (Phone / Push) | 5 minutes | Mandatory Runbook URL |
| **P2 - Warning** | Elevated latency, degraded redundancy, or single-AZ failure without customer loss | Slack `#alerts-payments` + Jira | 4 hours (business) | Dashboard Link |
| **P3 - Info** | Capacity trend, impending certificate renewal (> 30 days) | Datadog Event Stream | No active response | None |
#### 2. Symptom-Based Alert Triggers & PromQL Formulas [MC-AT-01]
##### Alert 1: PaymentTransactionErrorBurnRateCritical (P1)
Fires when error budget burns at 14.4x over 1 hour AND 14.4x over 5 minutes:
```promql
(
sum(rate(payment_requests_total{status="500"}[1h])) / sum(rate(payment_requests_total[1h])) > (14.4 * (1 - 0.999))
)
and
(
sum(rate(payment_requests_total{status="500"}[5m])) / sum(rate(payment_requests_total[5m])) > (14.4 * (1 - 0.999))
)
- Annotations:
summary: "High payment transaction error rate burning 2% of 30-day budget in 1 hour."runbook_url:https://runbooks.internal/payments/db-failover-remediation
Alert 2: PaymentAuthorizationLatencyP95Degraded (P1)
Fires when authorization latency p95 breaches 180 ms for 3 consecutive minutes:
histogram_quantile(0.95, sum(rate(payment_authorization_duration_seconds_bucket[3m])) by (le)) > 0.180
- Annotations:
runbook_url:https://runbooks.internal/payments/latency-triage-runbook
3. Deduplication, Grouping, and Inhibition Rules [MC-GI-01]
- Grouping Configuration:
group_by: ['alertname', 'cluster', 'service'] group_wait: 30s group_interval: 5m repeat_interval: 4h - Inhibition Hierarchy:
- If
PaymentClusterVPCUnreachableis firing:- Inhibit all
PaymentInstanceDownalerts. - Inhibit all
PaymentDatabaseConnectionTimeoutalerts.
- Inhibit all
- Prevents storm of hundreds of downstream alerts when the underlying root cause is a VPC gateway disconnect.
- If
4. On-Call Routing Matrix [MC-RM-01]
service: payment-core-> PagerDuty Schedulesched_pay_primary_sre.- Escalation: If unacknowledged after 10 minutes, escalate to Secondary SRE on-call; after 20 minutes, escalate to Marcus Vance (Lead SRE).
Invariants and Contracts
Mandatory Runbook URL Invariant [INV-ALT-01]
Every alert assigned severity `P1` must include an active, verified `runbook_url` in its annotations.
Alert definitions omitting runbook URLs are rejected by CI validation linters.
Zero Cause-Based Paging Invariant [INV-ALT-02]
Alerts triggering on raw infrastructure metrics (e.g. `node_cpu_utilization > 80%`, `jvm_memory_used > 80%`)
must never be configured with severity `P1`. Paging is strictly reserved for user-facing symptoms.
Deduplication Grouping Window Floor [INV-ALT-03]
Alertmanager configurations must enforce a `group_wait` of at least 30 seconds to allow related
component alerts to consolidate into a single notification batch before dispatch.
Explicit Unknowns
- Alertmanager webhook delivery retry count limits during external PagerDuty network outages (G-1).
- Effectiveness of Slack notification channel filtering during multi-service cloud provider zone brownouts (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Peak 1,400 transactions/sec | provided | Traffic profile intake | Current |
| Incident INC-3118 4,200 alert flood | provided | Post-mortem evidence | Historical |
| Prohibition of raw CPU alerts | decided | Marcus Vance (Lead SRE) | 2026-09-15 |
| Multi-window SLO burn rate (14.4x 1h / 5m) | decided | Google SRE Alerting Standard | 2026-09-15 |
| Mandatory runbook URL on P1 alerts | decided | Architectural invariant INV-ALT-01 | 2026-09-15 |
| Runbook URL reference | provided | https://runbooks.internal/payments/db-failover-remediation | Current |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against alerting design standards:
- Symptom Alignment: PASS. All P1 alerts track user transaction success or p95 latency.
- Noise Control: PASS. Grouping and inhibition rules collapse redundant pod failure notifications.
- Runbook Binding: PASS. Explicit
runbook_urlannotations provide direct remediation links. - Formatting Compliance: PASS. Conforms strictly to native Markdown rules in
rule_markdown.md.
Open Decisions
DEC-ALT-01: Marcus Vance to determine whether P2 warning alerts should post to a private Discord channel as a secondary notification channel (Owner: Marcus Vance).
Next steps
- Platform team validates Prometheus alert rules YAML using
promtool check rules. - Marcus Vance configures Alertmanager grouping and inhibition tree in
infra/monitoring/alertmanager.yml. - Conduct staging resilience drill triggering synthetic transaction failure to verify single grouped PagerDuty incident generation.
actionable-alert-contract-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted user-impact/objective conditions and telemetry into bounded detection, notification and response contracts. It defines when evidence should page, ticket or inform and how duplicate/noisy states are handled, without choosing SLOs or configuring a product.
Use it when
Use when one accepted failure/impact condition needs an exact signal, lifecycle, urgency, recipient and actionable response.
For example: “Our payment service drops transactions during database failovers, but engineers get paged 40 times for transient 1-second network blips. We need alert rules that only wake up on-call when real payment processing is blocked.”
What you get
- Alerting Policy Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/alerting-design/.
What it will not do
Do not use for defining SLOs/SLIs, dashboards, incident process, monitoring-product configuration or generic notification.
How it works
- Check alerting contract is required.
- Bind impact to explicit severity.
- Formulate condition query and evaluation clock.
- Define grouping and inhibition rules.
- Establish escalation routes and runbook links.
- Write the alerting policy spec under <output_root>/architecture/tasks/{run-id}/alerting-design/alerting-policy-spec.md.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install