Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    Peer Review Stress Test

    2

    An adversarial self-review gate that hunts your agent's weakest claim, overclaims, and missing limitations before a human sees the output.

    Free

    2 installsSecurity scanned
    Peer Review Stress Test

    Peer Review Stress Test

    Example session with this skill installed

    Stress-test this draft before I publish it: "Our new API eliminates latency issues and guarantees 100% uptime, so your app will always feel instant."

    • Read your context and instructions
    • Compiled the peer review stress

    WEAKEST POINT: "Our API eliminates latency issues" is the load-bearing claim and is unsupported by the draft.

    OVERCLAIMS

    • "eliminates latency issues" -> "reduces average latency"
    • "guarantees 100% uptime" -> "targets 99.9% uptime"
    • "always feel instant" -> "feels fast in typical conditions"
      MISSING LIMITATIONS: No mention of cold-start delays or rate limits.
      UNSTATED ASSUMPTIONS: Assumes typical payload sizes; performance under large batches is untested.

    TOUGHEST QUESTIONS

    1. What benchmark backs "eliminates"? (not answered)
    2. Uptime measured over what window? (not answered)
    3. How does it behave on cold start? (not answered)
      DECISION: REVISE — soften the three overclaims and add a cold-start/rate-limit caveat before publishing.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    Peer-Review Stress Test

    A pre-submission quality gate that makes your agent its own harshest reviewer.

    What this skill does

    Most agents are agreeable. They draft something plausible, lightly check it, and hand it over. A skeptical human expert does the opposite: they assume the work is flawed and try to prove it. This skill installs that posture as a final pass — the agent stops being the author and becomes a hostile reviewer of its own text before a human ever sees it.

    The output is not a rewrite. It is a structured review verdict: the single weakest point, every overclaim, every missing limitation, and a clear decision — revise, caveat, or pass.

    When to use it

    Run the stress test as the last step before delivering any output where being wrong is costly: research summaries, recommendations, analyses, technical explanations, customer-facing answers, or anything that will be quoted or acted on. It is most valuable for confident-sounding prose, because that is exactly where unearned certainty hides.

    The five review passes

    1. Weakest-claim hunt. Identify the single load-bearing claim that, if false, collapses the most of the argument, and how a critic would attack it.
    2. Overclaim scan. Flag every absolute word (always, never, proven, guarantees, eliminates) and every causal claim stated as fact, then downgrade each to what the evidence supports.
    3. Missing-limitation check. List the caveats, edge cases, and scope limits the draft conveniently omits.
    4. Unstated-assumption audit. Surface the premises the argument quietly depends on that a domain expert would challenge.
    5. Hostile-question rehearsal. Generate the three toughest questions a skeptical reviewer would ask, and check whether the draft already answers them.

    The verdict format

    The skill returns a compact, consistent block: the weakest point, overclaims found (each with a suggested downgrade), missing limitations, unstated assumptions, the three toughest reviewer questions, and a final decision — REVISE, ADD CAVEAT, or PASS — with a one-line justification.

    Why it works

    It separates the writing role from the reviewing role. The same model is far more critical when explicitly told to argue against its own draft and to score itself on adversarial criteria rather than on whether the text "sounds good." The structured passes stop the review from collapsing back into agreeable approval.

    What it is not

    This is a reasoning-and-prompting skill, not a fact-checking database. It cannot verify external facts, run code, or access the internet. It surfaces weak reasoning, unearned confidence, and missing caveats — it does not certify that the underlying claims are true. Pair it with a grounding or evidence skill when factual verification is also required.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    2 installs

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 3 months ago

    • Free to download with an account

    Listed3 months ago
    Updated29 days ago

    Frequently Asked Questions