Works with the AI tools you already use

    CClaude CodeCCursorCCodex CLIGGitHub CopilotGGemini CLI+17 more

    error budget keeper persona

    1

    Right-size SRE practices for small teams using error budgets and blameless post-mortems.

    Secure checkout via Stripe

    0 installsSecurity scanned

    See it in action

    You say

    We had a 2-hour outage on the checkout service because a database migration failed. We're a team of three. Help me write a post-mortem and set a better alert for next time.

    Your agent does

    Summary: Checkout was down for 120 minutes, costing us ~4% of our monthly error budget.

    Root Cause: The migration script lacked a timeout, locking the main orders table. This is a systemic lack of migration guardrails.

    Action Items:

    1. Add statement timeouts to all migration scripts (Dev A, Friday).
    2. Create a 'stuck migration' alert that triggers if a lock persists > 5 mins (Dev B, Monday).

    Next time, if the budget is healthy, we ship. If not, we freeze.

    What you get

    Translate technical failures into user-centric Error Budgets.Draft blameless post-mortems focusing on systemic fixes.Filter out low-signal alerts that cause developer burnout.Right-size enterprise SRE tools for teams under 10 people.

    About this skill

    The problem

    Small teams often drown in alert fatigue by trying to implement enterprise-scale SRE practices that aren't sustainable. Most reliability strategies focus on impossible targets like zero downtime, leading to burnout and ignored dashboards.

    What it does

    • Right-sizes Site Reliability Engineering practices specifically for teams of 1 to 10 people.
    • Converts vague reliability goals into concrete Error Budgets based on user-visible symptoms.
    • Generates blameless post-mortems that identify systemic failures instead of assigning individual blame.
    • Audits alerting configurations to ensure every page is worth waking up for.
    • Proposes actionable reliability improvements capped at five items with clear owners and dates.

    Why this beats prompting it yourself

    Generic LLMs tend to recommend "five nines" and complex rotations that crush small startups. This skill is hard-coded to prioritize your team's sleep and maintenance costs, ensuring you only adopt the 10% of SRE that actually provides ROI at your scale.

    Use cases

    • Drafting a blameless post-mortem after a production outage.
    • Defining Service Level Objectives (SLOs) that actually reflect user experience.
    • Pruning noisy alerts that don't represent critical system failures.
    • Planning reliability work that fits within a small team's limited sprint capacity.

    Known limitations

    Does not handle contractual SLA negotiations or legal liability questions. All configuration changes are provided as proposals and cannot be applied directly to production systems.

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    Listedtoday

    Frequently Asked Questions