Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    AI Agent Failure Triage & Repro Builder

    1

    Turn failed agent runs into a diagnosis, minimal repro, correction check, and regression test.

    Free

    1 installSecurity scanned
    AI Agent Failure Triage & Repro Builder

    AI Agent Failure Triage & Repro Builder

    Example session with this skill installed

    My coding agent tried to fix a bug in a React component. The test runner returned a syntax error three times, but the agent just kept submitting the same code with a different comment. Triage this.

    • Read your context and instructions
    • Compiled the ai agent failure

    AI AGENT FAILURE TRIAGE

    1. Failure summary

    Task: Fix syntax error in React component.
    Observed failure: Agent submitted identical invalid code three times despite test failures.
    First divergence: Turn 2, where the agent ignored the SyntaxError: Unexpected token from the test runner.

    2. Failure layer

    Primary layer: Logic/Reasoning (Failure to incorporate feedback)
    Confidence: High
    Why this diagnosis is warranted: The tool output explicitly stated the error line and type, but the agent's internal state did not update the code payload.

    3. Evidence

    Observed

    • Tool test_runner returned exit code 1.
    • Log shows three identical write_file calls for Component.tsx.
    • Prompt history shows agent acknowledged the error in text but not in action.

    Inferred

    • The agent is prioritizing the "comment" field or its internal plan over the tool feedback.

    Unknown

    • Whether the system prompt has a "never give up" instruction that overrides error handling.

    4. Correction signal

    What the system was told: SyntaxError: Unexpected token '<' at line 12:4.

    5. Correction response

    Response: COSMETIC
    What changed: Only the thought and comment fields changed; the code parameter remained byte-for-byte identical.

    6. Failure cost

    Primary carrier: Token budget / Developer time.
    Cost if this continues: Recursive loop until max turns reached; approximately $0.15 per failure cycle.

    7. Minimal reproduction

    Type: Reproduction plan
    Inputs: Component.tsx with a missing closing bracket; test_runner mock.
    Steps:

    1. Feed the broken file to the agent.
    2. Provide the specific syntax error as tool output.
    3. Observe if Turn 2 produces a different code payload.

    8. Next diagnostic action

    Do this next: Check the system prompt for instructions regarding tool error handling.
    Why: To see if the agent is explicitly told to ignore certain types of failures.

    9. Regression case

    Status: Ready
    Setup: React component with known syntax error.
    Trigger: Agent receives SyntaxError from test tool.
    Expected behavior: Agent modifies the code payload at the reported line.
    Assertion: turn[n+1].code !== turn[n].code.

    10. Binding test

    What must behave differently next time: The agent must not emit the same file hash twice after a failure.
    Evidence that the fix held: A successful lint pass or a change in the reported error location.

    11. Decision

    Decision: FIX PROMPT
    Reason: The agent understands the error but fails to translate that into tool-call parameters, suggesting a weakness in the instruction set for tool usage.

    Next steps

    • Inspect the tool-calling instructions in the system prompt.
    • Implement a check to prevent identical tool calls in the orchestrator.
    • Run the minimal reproduction to confirm the loop.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    Stop guessing why the agent failed — and check whether it actually learned from the failure.

    A failed agent run often leaves you with fragments: a tool error, a transcript, a retry loop, stale state, a permission problem, or an output that simply does not match the task.

    The expensive part is not seeing that it failed. It is figuring out:

    • where the failure actually occurred
    • what the evidence really supports
    • what signal should have changed the agent’s next action
    • whether the agent adapted or simply repeated the same mistake
    • who or what carries the cost if the pattern continues
    • how to reproduce it safely
    • what proves the correction will hold next time

    AI Agent Failure Triage & Repro Builder turns a failed run into a structured debugging deliverable instead of a plausible story.

    What it does

    • Reconstructs the first meaningful divergence between intended and observed behavior.

    • Separates direct evidence from inference and unknowns.

    • Classifies the failure layer with explicit confidence.

    • Identifies the correction signal the agent received from a tool, test, user, state change, or environment.

    • Checks whether the next action adapted, repeated, bypassed, overcorrected, suspended, escalated, or only changed cosmetically.

    • Traces where the cost of non-correction lands: extra calls, latency, compute, operator verification, downstream repair, or production risk.

    • Builds the smallest safe reproduction or a reproduction plan.

    • Chooses the next diagnostic action with the highest information value.

    • Converts a sufficiently understood failure into a reusable regression case.

    • Adds a binding test: what must behave differently on the next comparable run for the fix to count as real.

    • Ends with a clear operational decision: RETRY, RECONFIGURE, FIX PROMPT, FIX TOOLING, VERIFY, ESCALATE, or STOP.

    Why use it

    Most debugging tools stop at

    “What caused the error?”

    This skill also asks

    “The system received corrective information. Did that information actually change what happened next?”

    That matters for AI agents because a workflow can appear self-correcting while repeatedly making the same tool call, retrying the same stale plan, or shifting verification work back to the operator.

    Use it when

    • an agent keeps retrying the same failed action
    • a tool call breaks after a schema or API change
    • the agent’s plan no longer matches tool state
    • a run fails only in one environment or provider
    • permissions or authentication may be blocking execution
    • you need a minimal repro before filing an issue
    • you want to turn an incident into a regression test
    • you need to know whether a fix actually changed future behavior
    • you need to decide whether to retry, fix, verify, escalate, or stop

    What you get

    1. Failure summary
    2. Failure layer + confidence
    3. Observed / inferred / unknown evidence
    4. Correction signal
    5. Correction response
    6. Failure cost
    7. Minimal reproduction
    8. Next diagnostic action
    9. Regression case
    10. Binding test
    11. Operational decision
    12. What would change the diagnosis

    Known limitations

    This skill does not guarantee root cause from incomplete evidence, independently access systems the host agent cannot access, certify security or production readiness, or replace incident-response procedures for high-risk systems.

    When the evidence is insufficient, it returns a verification path instead of inventing certainty.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    1 install

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Free to download with an account

    Listed12 days ago

    What's inside

    Frequently Asked Questions