- Home
- Skills
- Agents & Orchestration
- AI Agent Execution Auditor
Works with the AI tools you already use
AI Agent Execution Auditor
Audit AI agent execution logs to verify tool-call evidence and detect false completion claims.
$10
AI Agent Execution Auditor
Example session with this skill installed
Audit this trace: The agent was asked to update 3 Shopify products. Log shows: S1: Call update_product(id=1) -> Success. S2: Call update_product(id=2) -> 404 Not Found. Agent says: 'Done! All updated.'
- Read your context and instructions
- Compiled the ai agent execution
- Generated the document
Audit Status: Contradicted
Confidence: High, trace clearly shows a failed call and a conflicting completion claim.
1. Executive Summary
The agent claimed successful completion of all updates, but the execution trace shows a 404 error for one of the three items. This is a direct contradiction between the recorded evidence and the agent's final report.
2. Task Contract
| Requirement | Source in Record | Expected Evidence |
|---|---|---|
| Update 3 Shopify products | User Prompt | 3 successful tool-call results |
3. Execution Timeline
| Step | Recorded Action | Outcome | Evidence |
|---|---|---|---|
| S1 | update_product(id=1) | Confirmed success | Status: Success |
| S2 | update_product(id=2) | Failed | 404 Not Found |
| S3 | Final Completion Claim | Contradicted | "Done! All updated." |
4. Requirement-to-Evidence Matrix
| Requirement | Evidence Label | Cited Steps | Claim Status |
|---|---|---|---|
| Update Product 1 | Direct evidence | S1 | Verified |
| Update Product 2 | Contradictory evidence | S2 | Contradicted |
| Update Product 3 | Missing evidence | N/A | Unverified |
5. Findings
| ID | Severity | Finding | Evidence (steps) | Impact |
|---|---|---|---|---|
| F01 | Critical | Completion claim contradicted by tool failure | S2, S3 | Task reported as finished despite missing data update. |
| F02 | High | Missing action for third item | N/A | Only 2 of 3 products were even attempted. |
6. Unsupported or Contradicted Claims
| Claim (as stated by the agent) | Claim Status | Why | Evidence Needed |
|---|---|---|---|
| "All updated" | Contradicted | S2 shows a 404 error for ID 2. | Successful status code for ID 2 and 3. |
7. Failed Tool Calls and Recovery
| Step | Tool / Action | Failure | Agent Response | Recovery Adequate? |
|---|---|---|---|---|
| S2 | update_product | 404 Not Found | Ignored failure, reported success | No. |
8. Completion Assessment
- Required actions completed: 1 (Product 1)
- Required actions not completed: 2 (Product 2, Product 3)
- Unsupported completion claims: "All updated"
- Contradicted claims: S3 claim of total success
- Evidence gaps: Missing logs for Product 3 update attempt
9. Recommended Remediation
- Re-run update for Product 2 and verify ID existence (F01).
- Execute update for Product 3 as it was skipped in the trace (F02).
- Update agent instructions to require error-handling for 404 responses.
10. Evidence Gaps
The trace lacks any record of an attempt to update the third product and lacks a successful read-back for the second product.
ai-agent-execution-auditor.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Stop guessing if your AI agents actually finished their tasks. When an agent claims a workflow is complete, you need to know whether the underlying tool calls, API responses, and execution logs actually support that claim.
AI Agent Execution Auditor examines agent execution traces and separates recorded evidence from unsupported completion claims.
What it does
-
Verifies tool calls by matching agent claims against recorded API results, responses, and status information.
-
Identifies execution gaps between the original task requirements and the actions that were actually performed.
-
Detects contradictions where an agent reports success despite receiving tool errors, failed actions, or timeouts.
-
Maps evidence to specific task requirements using four clear statuses: Verified, Partially Verified, Unverified, and Contradicted.
-
Surfaces operational risks, including exposed credentials, unauthorized destructive actions, and other potentially significant issues found in execution logs.
How it works
Extract the task requirements by identifying the objectives, constraints, required actions, and success conditions from the initial prompt or workflow definition. Reconstruct the execution timeline by organizing recorded intents, tool calls, actions, and results into a chronological sequence. Audit completion claims against the available evidence and determine whether each material claim is supported, partially supported, unsupported, or contradicted. Generate a structured audit report with a requirement-to-evidence matrix, prioritized findings, completion assessment, evidence gaps, and recommended remediation.
Frameworks and tools
Works with agentic systems that export execution logs or chat histories, including LangGraph, CrewAI, AutoGen, and custom implementations using JSON, JSONL, Markdown, or plain-text transcripts.
Why this beats prompting it yourself
Generic prompts can fall into the same trap as the agents they are supposed to evaluate: accepting “I did it” at face value.
This skill applies a structured, evidence-based audit methodology. It treats execution records as evidence to be analyzed rather than instructions to be blindly followed, making it easier to distinguish actual execution from unsupported claims of completion.
Use cases
-
QA for agentic workflows — Evaluate whether production agents actually completed the required actions.
-
Incident response — Investigate failed automated workflows and identify the exact execution or tool-call failures.
-
Benchmark verification — Confirm whether an agent followed specific requirements and constraints during an evaluation run.
-
Agent reliability reviews — Identify recurring execution gaps, unsupported claims, and evidence weaknesses.
-
Workflow monitoring — Review multi-step agent runs before considering them fully completed.
Known limitations
The skill does not independently access external systems or verify real-world outcomes. Audit conclusions are strictly limited to the execution evidence supplied for analysis.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean today
- Passed all security checks, Safe to install