Agent Optimization & Output-Quality Suite
A tight loop for making AI agent output measurably better and proving it. Agent Loop iteratively optimizes any prompt, config, or artifact - changing one thing at a time, scoring it against your metric, and keeping only the winners. The AI Eval & Test-Suite Quality Gate makes sure the metric you optimize against is trustworthy, catching gameable criteria, data leakage, and missing edge cases before they mislead you. And the Peer-Review Stress Test turns your agent into its own harshest reviewer, hunting weak claims and missing limitations before a human sees the output. Built for prompt engineers, AI builders, and teams shipping agent-powered features. Optimize it, trust the score, then stress-test it.
Works with every agent that reads SKILL.md — Claude Code, Cursor, Codex CLI, Gemini CLI, GitHub Copilot, Windsurf, OpenClaw, and more.
One payment, lifetime access. 3 skills unlock instantly in your library.
30-day refund guarantee
Instant unlock in your library
Free updates from the creator
What's included
3 skills
An adversarial gate that audits an AI eval or test suite — LLM-judge rubrics, datasets, regression tests, metrics — for gameable criteria, data leakage, missing edge cases, and non-determinism, then returns one PASS/REVISE/FAIL verdict.

An adversarial self-review gate that hunts your agent's weakest claim, overclaims, and missing limitations before a human sees the output.

An iterative agent loop that optimizes any prompt, config, or artifact by making one change at a time, scoring it against a metric, and keeping only the winners.
See it in action
You say
Audit this prompt-regression suite for a support agent. Grader prompt: "Rate the answer 1-5 on helpfulness and completeness. A good answer is detailed and covers the topic." 40 test cases, each with its golden answer included in the same prompt block shown to the model. Reported result: average 4.6/5, ship it.
Your agent does
EVAL TRUST: FAIL WEAKNESSES FOUND:
- [gameable-rubric] "detailed and covers the topic" rewards length and breadth, not factual correctness -> a verbose wrong answer scores high
- [data-leakage] golden answers sit in the same block sent to the model -> the model can copy them; 4.6/5 measures copying, not capability
- [determinism] single-run average, no seed, temperature, or variance -> 4.6 is one sample, not a stable measure
- [metric-validity] one averaged helpfulness number hides whether any answer was actually wrong BLIND SPOT: A regression that makes the agent confidently incorrect would still average near 4.6 and pass. VERDICT: FAIL - remove golden answers from the model's context, anchor the rubric to correctness, and report seeded multi-run results before trusting any score.
How to install
Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.