BUNDLE Security scanned3 skills

    Agent Optimization & Output-Quality Suite

    A tight loop for making AI agent output measurably better and proving it. Agent Loop iteratively optimizes any prompt, config, or artifact - changing one thing at a time, scoring it against your metric, and keeping only the winners. The AI Eval & Test-Suite Quality Gate makes sure the metric you optimize against is trustworthy, catching gameable criteria, data leakage, and missing edge cases before they mislead you. And the Peer-Review Stress Test turns your agent into its own harshest reviewer, hunting weak claims and missing limitations before a human sees the output. Built for prompt engineers, AI builders, and teams shipping agent-powered features. Optimize it, trust the score, then stress-test it.

    365 views

    Works with every agent that reads SKILL.md — Claude Code, Cursor, Codex CLI, Gemini CLI, GitHub Copilot, Windsurf, OpenClaw, and more.

    PubsProToolkit
    Created by
    PubsProToolkit
    $24$31
    Save 23% · $7

    One payment, lifetime access. 3 skills unlock instantly in your library.

    30-day refund guarantee

    Instant unlock in your library

    Free updates from the creator

    What's included

    3 skills

    See it in action

    You say

    Audit this prompt-regression suite for a support agent. Grader prompt: "Rate the answer 1-5 on helpfulness and completeness. A good answer is detailed and covers the topic." 40 test cases, each with its golden answer included in the same prompt block shown to the model. Reported result: average 4.6/5, ship it.

    Your agent does

    EVAL TRUST: FAIL WEAKNESSES FOUND:

    • [gameable-rubric] "detailed and covers the topic" rewards length and breadth, not factual correctness -> a verbose wrong answer scores high
    • [data-leakage] golden answers sit in the same block sent to the model -> the model can copy them; 4.6/5 measures copying, not capability
    • [determinism] single-run average, no seed, temperature, or variance -> 4.6 is one sample, not a stable measure
    • [metric-validity] one averaged helpfulness number hides whether any answer was actually wrong BLIND SPOT: A regression that makes the agent confidently incorrect would still average near 4.6 and pass. VERDICT: FAIL - remove golden answers from the model's context, anchor the rubric to correctness, and report seeded multi-run results before trusting any score.

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet on the included skills. Be the first to try this bundle.

    Frequently asked questions

    More bundles from PubsProToolkit