More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    Agent pr Verifier

    1

    Your agent spent 40 minutes on a feature and closed with "All 23 tests passing, lint clean, PR ready." You look at the diff — it looks right. You merg

    Free

    1 installSecurity scanned
    agent-pr-verifier

    agent-pr-verifier

    Example session with this skill installed

    ** "Agent just finished the refund-retry feature and says all tests pass. Verify before I merge PR #482.

    • Read your context and instructions
    • Compiled the agent-pr-verifier
    • Generated the document

    **

    EVIDENCE REPORT — PR #482 (refund-retry)
    Claims found: 4 | Re-runnable: 3 | Unverifiable: 1
    
    === VERIFIED (re-ran from clean state, HEAD a91c3f2) ===
      pytest tests/ -q                    → 23 passed (exit 0)
      ruff check .                        → 0 findings (exit 0)
      npm run build                       → success (exit 0)
    
    === VACUOUS (checks that cannot fail) ===
      test_refund_retry_calls_gateway
        Deletion probe: stubbed RefundGateway.send() → test STILL PASSES
        Reason: asserts on mock return value, never touches implementation
        Proof probe: git stash && pytest tests/test_refund.py::test_refund_retry_calls_gateway
        → 1 of 23 tests verified nothing
    
    === GAPS (requirement clauses with no test) ===
      Requirement: "retry with exponential backoff, max 3 attempts"
        Tests exercise retry ONCE. No test covers attempt 2 or 3, or the
        give-up path after max retries.
      R
    

    agent-pr-verifier.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    "Tests pass" from an AI agent is a report of intent, not an observation. This skill turns the reviewer's side into a mechanical process — and finds the checks that cannot fail.

    The Problem

    Your agent spent 40 minutes on a feature and closed with "All 23 tests passing, lint clean, PR ready." You look at the diff — it looks right. You merge. Three days later production breaks in exactly the way the tests were supposedly covering. The post-mortem finds the trap: the agent wrote the implementation AND the tests from the same wrong reading of the requirement, so they agree perfectly — on the wrong behavior. Or worse: the "coverage" test passes whether or not the feature exists, because it asserts a mock, not the code. You had 23 green checks and zero evidence.

    What You Get

    • Claim inventory extraction — pulls every checkable claim from the agent's final report ("tests pass", "lint clean", "verified in browser") and classifies each as re-runnable, stale, or unverifiable.
    • Independent re-run — executes each claim's check from a clean state against current HEAD, with exact commands and exit codes recorded — never the agent's terminal history.
    • The deletion probe — temporarily deletes or stubs the code each test claims to cover and confirms the test FAILS. A test that still passes when the feature is deleted is vacuous, and gets flagged with a one-command proof probe.
    • The zero-findings probe — plants a known-bad artifact into scanners that reported "0 issues" and confirms the scanner catches it. A scanner with a matching problem reads as "clean" forever.
    • Shared-misunderstanding detection — flags tests written in the same PR as the implementation whose assertions restate the implementation instead of the original requirement, and requirement clauses no test exercises.
    • One-screen evidence report — VERIFIED / VACUOUS / GAPS sections with a merge recommendation (SAFE / CONDITIONAL / BLOCKED) and the single next command that provides the strongest remaining evidence.

    Why This Beats Prompting It Yourself

    Ask your agent to "double-check your work" and it re-runs its own test command and says "all good" — the same instrument that reported success the first time, re-run by the party with the incentive to see green. It will never delete the code under test to see if the test notices, because that requires suspecting its own work. This skill mechanically distrusts the report: re-runs from clean state, probes for vacuity, and compares tests against the requirement text instead of the implementation. It is the reviewer-side complement to agent self-verification — designed for the case where the "author" is a session that no longer exists.

    Example Session

    You: "Agent just finished the refund-retry feature and says all tests pass. Verify before I merge PR #482."

    Agent:

    EVIDENCE REPORT — PR #482 (refund-retry)
    Claims found: 4 | Re-runnable: 3 | Unverifiable: 1
    
    === VERIFIED (re-ran from clean state, HEAD a91c3f2) ===
      pytest tests/ -q                    → 23 passed (exit 0)
      ruff check .                        → 0 findings (exit 0)
      npm run build                       → success (exit 0)
    
    === VACUOUS (checks that cannot fail) ===
      test_refund_retry_calls_gateway
        Deletion probe: stubbed RefundGateway.send() → test STILL PASSES
        Reason: asserts on mock return value, never touches implementation
        Proof probe: git stash && pytest tests/test_refund.py::test_refund_retry_calls_gateway
        → 1 of 23 tests verified nothing
    
    === GAPS (requirement clauses with no test) ===
      Requirement: "retry with exponential backoff, max 3 attempts"
        Tests exercise retry ONCE. No test covers attempt 2 or 3, or the
        give-up path after max retries.
      Requirement: "idempotent by event ID"
        No test submits the same event ID twice.
      Claim "no breaking changes": UNVERIFIABLE — no API contract check exists
        Strongest next command: npm run api-diff --Against=main
    
    === SHARED-MISUNDERSTANDING CHECK ===
      test_refund_amount_matches_request
        Asserts: refund.amount == request.amount
        Requirement says: "refund amount EXCLUDING original transaction fee"
        Implementation and test agree with each other — and both disagree
        with the requirement. HIGH severity.
    
    MERGE RECOMMENDATION: BLOCKED
      1 vacuous test, 2 untested requirement clauses, 1 requirement divergence
      Fix path: assert on gateway side effects (not mock), add attempt-3 and
      duplicate-event-ID tests, resolve fee-handling against the requirement.
    

    Use Cases

    • Verifying an agent-authored feature PR before merging to a shared branch.
    • Orchestrator sessions validating executor session claims ("tests pass" → prove it).
    • Auditing a large agent-generated codebase to find which parts have real evidence and which are Ferrari shells.
    • CI augmentation: run the deletion probe on the 10 most critical tests as a scheduled job.
    • Solo developers using agents as their only implementer — this is the second pair of eyes.

    Known Limitations

    The deletion probe mutates the working tree and must run on stashed/clean state — it is not safe on uncommitted work-in-progress. It verifies evidence quality, not code style or architecture. Detecting shared misunderstanding requires the original requirement text to be available and specific; vague tickets weaken the check. Runtime/concurrency bugs that no static probe exposes remain the reviewer's judgment call.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    1 install

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 11 days ago

    • Free to download with an account

    Listed11 days ago
    Updated9 days ago

    What's inside

    Frequently Asked Questions