Works with the AI tools you already use
skill regression harness
Generate a cost-tiered regression testing harness to catch silent behavioral breakages after editing SKILL.md files.
$16
skill regression harness
Example session with this skill installed
I just updated the Iron Rules in skills/billing-agent.md to be stricter about data privacy. Generate a regression harness to make sure I didn't break the existing refund trigger.
- Read your context and instructions
- Compiled the regression harness
I have analyzed skills/billing-agent.md and designed a 35-case regression battery. This harness is configured to use Claude Haiku 4.5 for bulk testing and Claude Opus 5 for escalation and confirmation sampling.
Regression Harness for Billing Agent
Estimated Run Cost:
- Bulk Pass (Haiku 4.5): $0.07
- Judging & Escalation (Opus 5): $0.12
- Total Projected Cost: ~$0.19 per run
The following files have been generated in tests/regression/billing-agent/:
config.yaml: Defines model routing, API environment variables, and a$0.50safety ceiling.battery.yaml: 15 trigger-positive cases (refunds), 10 trigger-negative cases (general billing), and 10 compliance cases mapped to your new privacy rules.run_battery.py: Script to generate responses from the skill under test.score_responses.py: Rubric-based grader that performs cost-tiered escalation.README.md: Setup instructions and usage guide.
# config.yaml snippet
models:
bulk: "claude-3-haiku-20240307"
frontier: "claude-3-opus-20240229"
routing:
confirmation_sample_rate: 0.1
escalate_all_failures: true
costs:
max_projected_cost_usd: 0.50
Next steps
- Review the test cases in
battery.yamlto ensure they reflect your expected "strict" privacy behaviors. - Set your
ANTHROPIC_API_KEYenvironment variable. - Run
python run_battery.py --estimate-onlyto verify the setup before making live API calls.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Editing a shipped SKILL.md is like editing a live program without a compiler. A single wording change can silently break trigger accuracy or cause the agent to ignore its own Iron Rules. This skill solves the problem of expensive, manual re-testing by generating a portable, automated regression harness for your specific skill files. It provides the scripts you need to catch behavior drift before merging edits, using a cost-efficient model routing strategy that reserves frontier models only for critical verifications.
What it does
- Battery generation creates a structured set of 25 to 40 test cases covering triggers, compliance, and adversarial inputs.
- Cost-tiered routing runs bulk tests on cheaper models like Haiku 4.5 and escalates failures to Opus 5 for confirmation.
- Rubric-based scoring extracts specific rules from your skill text to judge outputs with traceable citations instead of vague scores.
- Regression triage compares failures against your recent edits to distinguish between silent breakages and intended behavior changes.
- Budget enforcement calculates and enforces API cost ceilings for both response generation and grading passes.
How it works
- Analyze target skill to identify all trigger surfaces and compliance rules within your SKILL.md.
- Design prompt battery including positive/negative triggers and rubric-linked adversarial prompts.
- Generate harness code including
config.yaml, test scripts, and model routing logic for your local environment. - Execute local scripts to run the battery, score responses, and generate a detailed
scorecard.mdreport.
Frameworks & tools
This skill generates Python-based scripts (run_battery.py, score_responses.py) that interface with the Anthropic API. It uses YAML for configuration and Markdown for reporting.
Why this beats prompting it yourself
Manual testing is inconsistent and expensive if you use frontier models for every check. This skill automates the "bulk-test, frontier-escalate" workflow, ensuring you only pay for high-end reasoning when a cheaper model flags a potential issue.
Use cases
- Validating a PR that modifies core agent instructions in a production skill.
- Setting up CI/CD gates to prevent behavioral regressions in a skill portfolio.
- Auditing third-party skills for rule compliance and adversarial robustness.
- Reducing API spend during the maintenance phase of an agent's lifecycle.
Known limitations
This tool is for post-ship regression testing, not interactive first-draft authoring. It requires local Python execution and valid Anthropic API credentials.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 1 day ago
- Passed all security checks, Safe to install
Needs access to