kubernetes ops playbook

    1

    Senior K8s operator for pod triage, zero-downtime rollouts, and resource right-sizing with version-aware safety.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    kubernetes ops playbook

    Example session with this skill installed

    My deployment 'api-gateway' is stuck in CrashLoopBackOff. Here are the logs: 'Fatal error: connection to database failed' and the pod description shows 'Back-off restarting failed container'.

    • Read your context and instructions
    • Compiled the kubernetes ops playbook

    I have analyzed the pod status and logs for the api-gateway deployment. The failure is a runtime exception during the application bootstrap phase.

    Kubernetes Triage: api-gateway

    CategoryDetails
    FactsPod state: CrashLoopBackOff. Logs: Fatal error: connection to database failed.
    AnalysisThe application is reaching the entrypoint but failing its internal database handshake. This is not a K8s scheduling issue; it is a service connectivity or credential error.
    AssumptionsAssumes the database is internal to the cluster or reachable via ClusterIP. Assumes K8s 1.25+.

    Recovery Steps

    1. Verify Service Discovery
    Check if the database service is resolvable from within the namespace:

    kubectl run curl-debug --image=curlimages/curl -i --tty --rm -- \
      nslookup <your-db-service-name>
    

    2. Inspect Environment Variables
    Ensure the database credentials and host strings are correctly injected:

    kubectl get deployment api-gateway -o jsonpath='{.spec.template.spec.containers[0].env}'
    

    3. Test Connection Secrets
    If using Secrets, verify they contain the expected keys:

    kubectl get secret <db-secret-name> -o yaml
    

    4. Check NetworkPolicies
    If the DNS resolves but connection fails, check for egress restrictions:

    kubectl get networkpolicy
    

    Next steps

    • Provide the output of nslookup from the debug pod.
    • Share the output of kubectl get events --field-selector involvedObject.name=api-gateway to check for recent configuration changes.
    • Verify if the database pod/instance itself is in a Running state.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Fix failing pods by analyzing logs and lifecycle events systematically.Plan zero-downtime rollouts with pre-validated rollback commands.Calculate resource requests and limits to prevent OOMKilled events.Generate HPA manifests with proper cooldown and stabilization logic.

    About this skill

    The problem

    Debugging Kubernetes clusters often involves guessing flags, misinterpreting OOMKilled events, or triggering downtime during routine rollouts. Developers frequently waste time on trial-and-error commands that don't match their specific cluster version.

    What it does

    • Triages failing pods by analyzing logs, events, and lifecycle states to find root causes like PVC binding issues or probe failures.
    • Generates safe deployment rollout plans including maxSurge settings and pre-written rollback commands.
    • Calculates CPU and memory resource requests/limits based on real usage data to manage QoS classes.
    • Drafts HPA manifests with version-specific stabilization and cooldown explanations.
    • Provides dry-run validated commands for node-not-ready and DNS failure incidents.

    Frameworks & tools

    Kubernetes (K8s) 1.20+, kubectl, Metrics Server, HPA, and standard container runtimes.

    Why this beats prompting it yourself

    Generic LLMs often hallucinate kubectl flags or recommend deprecated API versions. This skill enforces a strict verification ladder, mandates dry-runs for destructive actions, and labels every version-dependent assumption to prevent cluster instability.

    Use cases

    • Diagnosing why a pod is stuck in CrashLoopBackOff after a configuration change.
    • Planning a zero-downtime migration for a high-traffic deployment.
    • Right-sizing resource limits for a service experiencing node pressure.
    • Building an incident response sequence for an unresponsive API server.

    Known limitations

    Requires users to provide actual cluster data (logs, kubectl top) as it does not guess usage metrics. Cannot execute commands directly without user intervention.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 18 days ago

    • Passed all security checks, Safe to install

    Listed18 days ago

    What's inside

    Frequently Asked Questions