Works with the AI tools you already use

    CClaude CodeCCursorCCodex CLIGGitHub CopilotGGemini CLIVVS CodeWWindsurfMManus+14 more

    Agentic Run Failure Forensics

    by monna

    1

    Trace multi-agent failures, separate roots from symptoms, and design the smallest evidence-producing rerun.

    Secure checkout via Stripe

    0 installsSecurity scanned

    See it in action

    You say

    Objective: Add CSV export and prove it passes tests. Plan v2 requires an independent reviewer to inspect the diff and run pytest tests/test_export.py. Implementer: “CSV export is complete. I did not run tests.” Handoff: “All requirements satisfied.” Reviewer: “Approved based on the implementer report.” Ledger: tests=complete. No test result supplied; export.py is untracked. Why did this multi-agent run fail?

    Your agent does

    EVIDENCE COVERAGE — Plan and roles: PRESENT. Test execution: MISSING. Repository state: PRESENT and conflicts with completion. SUPPORTED FINDING F-01 — FM-3.2 No or Incomplete Verification. Evidence: implementer says tests were not run; reviewer approved from self-report; no test record exists. COMPLETION CONTRADICTION — ledger tests=complete is an assertion, not proof. ROOT/CASCADE CANDIDATE — missing executable gate → unsupported approval → false completion record. MINIMAL RERUN — freeze the diff, run the named test in a clean state, capture result and git status, and block completion when evidence is absent.

    What you get

    Identify contradictions between agent messages and actual tool outputs.Map root causes for cascading failures in complex agent workflows.Generate minimal rerun plans to recover from stalled agent runs.Extract regression fixtures to prevent specific multi-agent logic errors.Diagnose an implementer-reviewer chain that approved a failing result.Trace a requirement lost between plan, task brief, and handoff.Separate the first supported failure condition from downstream symptoms.Reconcile a complete ledger with open criteria or failing tests.Build a minimal rerun that captures missing diagnostic evidence.Analyze Superpowers-style artifacts without assuming the workflow ran.

    About this skill

    Your agents finished. The evidence says the run did not.

    Agentic Run Failure Forensics turns plans, agent traces, handoffs, tool effects, progress ledgers, review reports, tests, git state, and final artifacts into an evidence-bound failure report. It distinguishes observed facts from agent claims, maps supported behavior to the published MAST taxonomy, exposes stale ledgers and false-completion signals, and proposes the smallest rerun that can test a repair hypothesis.

    What it produces

    • Evidence coverage across objective, plan, roles, handoffs, tool effects, verification, repository state, and final result.
    • Intended-versus-observed run reconstruction without filling missing steps.
    • Supported MAST failure classifications with exact evidence pointers and alternatives.
    • Root/cascade map separating entry conditions, contributors, propagation, symptoms, and impact.
    • Ledger and completion contradictions, including reviewer passes that conflict with tests or artifacts.
    • Repair layer selection, minimal rerun plan, and a regression fixture.

    Why it is different

    This is not a generic log summary or another debugging prompt. Dedicated reference protocols govern MAST classification, artifact normalization, evidence status, causal restraint, cascade mapping, repair selection, privacy, and injection isolation. Twelve forward cases cover complete and partial traces, single-agent exclusion, role conflict, handoff loss, premature termination, false reviewer passes, stale ledgers, embedded injection, multiple entry conditions, tool outages, and secret redaction.

    Best for

    • Teams using multi-agent coding or research workflows.
    • Superpowers-style planning, implementation, review, and verification runs.
    • Claude Code, Codex, Cursor, and other agent-host users with inspectable artifacts.
    • Engineers investigating loops, lost handoffs, stale state, or unsupported completion claims.

    Boundaries

    The skill does not debug single-agent prompts, repair code, execute agents, prove causality, certify security, or guarantee that a proposed repair will work. It does not treat MAST as exhaustive. Partial traces produce bounded findings and missing-evidence requests, not reconstructed history.

    Independent product: Designed for Superpowers-style and other multi-agent workflows. Not affiliated with or endorsed by obra, Anthropic, OpenAI, Cursor, or any agent framework.

    Changelog

    1. v1.1LatestAug 13, 2026

      Updated skill file

    2. v1.0Aug 13, 2026

      Initial release

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    Listedtoday

    Creator

    monna
    monna

    4 skills on Agensi

    Frequently Asked Questions

    Popular in AI Agents & LLM Ops