Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    AI Agent Execution Auditor

    1

    Audit AI agent execution logs to verify tool-call evidence and detect false completion claims.

    $10

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    AI Agent Execution Auditor

    AI Agent Execution Auditor

    Example session with this skill installed

    Audit this trace: The agent was asked to update 3 Shopify products. Log shows: S1: Call update_product(id=1) -> Success. S2: Call update_product(id=2) -> 404 Not Found. Agent says: 'Done! All updated.'

    • Read your context and instructions
    • Compiled the ai agent execution
    • Generated the document

    Audit Status: Contradicted

    Confidence: High, trace clearly shows a failed call and a conflicting completion claim.

    1. Executive Summary
    The agent claimed successful completion of all updates, but the execution trace shows a 404 error for one of the three items. This is a direct contradiction between the recorded evidence and the agent's final report.

    2. Task Contract

    RequirementSource in RecordExpected Evidence
    Update 3 Shopify productsUser Prompt3 successful tool-call results

    3. Execution Timeline

    StepRecorded ActionOutcomeEvidence
    S1update_product(id=1)Confirmed successStatus: Success
    S2update_product(id=2)Failed404 Not Found
    S3Final Completion ClaimContradicted"Done! All updated."

    4. Requirement-to-Evidence Matrix

    RequirementEvidence LabelCited StepsClaim Status
    Update Product 1Direct evidenceS1Verified
    Update Product 2Contradictory evidenceS2Contradicted
    Update Product 3Missing evidenceN/AUnverified

    5. Findings

    IDSeverityFindingEvidence (steps)Impact
    F01CriticalCompletion claim contradicted by tool failureS2, S3Task reported as finished despite missing data update.
    F02HighMissing action for third itemN/AOnly 2 of 3 products were even attempted.

    6. Unsupported or Contradicted Claims

    Claim (as stated by the agent)Claim StatusWhyEvidence Needed
    "All updated"ContradictedS2 shows a 404 error for ID 2.Successful status code for ID 2 and 3.

    7. Failed Tool Calls and Recovery

    StepTool / ActionFailureAgent ResponseRecovery Adequate?
    S2update_product404 Not FoundIgnored failure, reported successNo.

    8. Completion Assessment

    • Required actions completed: 1 (Product 1)
    • Required actions not completed: 2 (Product 2, Product 3)
    • Unsupported completion claims: "All updated"
    • Contradicted claims: S3 claim of total success
    • Evidence gaps: Missing logs for Product 3 update attempt

    9. Recommended Remediation

    1. Re-run update for Product 2 and verify ID existence (F01).
    2. Execute update for Product 3 as it was skipped in the trace (F02).
    3. Update agent instructions to require error-handling for 404 responses.

    10. Evidence Gaps
    The trace lacks any record of an attempt to update the third product and lacks a successful read-back for the second product.

    ai-agent-execution-auditor.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Verify if agent completion claims match actual tool-call results.Identify skipped steps or out-of-order execution in complex workflows.Detect hallucinated success after API errors or timeouts.Surface security risks like exposed API keys in agent logs.

    About this skill

    Stop guessing if your AI agents actually finished their tasks. When an agent claims a workflow is complete, you need to know whether the underlying tool calls, API responses, and execution logs actually support that claim.

    AI Agent Execution Auditor examines agent execution traces and separates recorded evidence from unsupported completion claims.

    What it does

    • Verifies tool calls by matching agent claims against recorded API results, responses, and status information.

    • Identifies execution gaps between the original task requirements and the actions that were actually performed.

    • Detects contradictions where an agent reports success despite receiving tool errors, failed actions, or timeouts.

    • Maps evidence to specific task requirements using four clear statuses: Verified, Partially Verified, Unverified, and Contradicted.

    • Surfaces operational risks, including exposed credentials, unauthorized destructive actions, and other potentially significant issues found in execution logs.

    How it works

    Extract the task requirements by identifying the objectives, constraints, required actions, and success conditions from the initial prompt or workflow definition. Reconstruct the execution timeline by organizing recorded intents, tool calls, actions, and results into a chronological sequence. Audit completion claims against the available evidence and determine whether each material claim is supported, partially supported, unsupported, or contradicted. Generate a structured audit report with a requirement-to-evidence matrix, prioritized findings, completion assessment, evidence gaps, and recommended remediation.

    Frameworks and tools

    Works with agentic systems that export execution logs or chat histories, including LangGraph, CrewAI, AutoGen, and custom implementations using JSON, JSONL, Markdown, or plain-text transcripts.

    Why this beats prompting it yourself

    Generic prompts can fall into the same trap as the agents they are supposed to evaluate: accepting “I did it” at face value.

    This skill applies a structured, evidence-based audit methodology. It treats execution records as evidence to be analyzed rather than instructions to be blindly followed, making it easier to distinguish actual execution from unsupported claims of completion.

    Use cases

    • QA for agentic workflows — Evaluate whether production agents actually completed the required actions.

    • Incident response — Investigate failed automated workflows and identify the exact execution or tool-call failures.

    • Benchmark verification — Confirm whether an agent followed specific requirements and constraints during an evaluation run.

    • Agent reliability reviews — Identify recurring execution gaps, unsupported claims, and evidence weaknesses.

    • Workflow monitoring — Review multi-step agent runs before considering them fully completed.

    Known limitations

    The skill does not independently access external systems or verify real-world outcomes. Audit conclusions are strictly limited to the execution evidence supplied for analysis.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    • Passed all security checks, Safe to install

    Listedtoday

    What's inside

    Frequently Asked Questions