Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    AI Agent Regression Testing & Release Gate.

    1

    Deterministic regression testing for AI prompts and agentic tool-use workflows using captured JSONL run records.

    $10

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    AI Agent Regression Testing & Release Gate.

    AI Agent Regression Testing & Release Gate.

    Example session with this skill installed

    Create a test suite for a weather agent that must call the get_weather tool with a city name and verify the final response contains a temperature.

    • Read your context and instructions
    • Compiled the ai agent regression

    I have generated a deterministic test suite for the weather agent. This suite validates both the tool execution and the presence of data in the final response.

    Weather Agent Test Suite

    {
      "suite_name": "Weather Agent Verification",
      "test_cases": [
        {
          "case_id": "weather-lookup-san-francisco",
          "title": "Basic Weather Query",
          "input": "What is the weather in San Francisco?",
          "assertions": [
            {
              "type": "tool_call_exists",
              "tool_name": "get_weather",
              "expected_arguments": { "city": "San Francisco" }
            },
            {
              "type": "text_contains",
              "pattern": "degrees"
            },
            {
              "type": "json_schema",
              "path": "metadata",
              "schema": { "type": "object", "required": ["model", "run_id"] }
            }
          ],
          "tags": ["smoke-test", "core-functionality"]
        }
      ]
    }
    

    Next steps

    • Capture a run of your weather agent into a run-record.jsonl file.
    • Execute the validation script: python scripts/prompt_suite.py weather_suite.json --run run-record.jsonl.
    • Review the generated eval.md report to confirm the tool arguments were mapped correctly by the model.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Validate tool call sequences and arguments against a known-good baseline.Enforce strict JSON schema compliance for prompt outputs in CI/CD.Run regression tests on captured model traces without incurring API costs.Tag and filter safety-critical test cases for specialized audit reports.

    About this skill

    Regression in AI behavior often goes unnoticed until a user reports it. Prompt changes can break JSON structures, skip tool calls, or fail safety checks without warning. This skill provides a deterministic framework for validating agentic workflows against captured run records.

    What it does

    • Deterministic validation compares captured agent outputs and tool traces against a set of predefined assertions.
    • Tool call auditing verifies that specific tools were invoked in the correct order with the expected arguments.
    • Regression monitoring flags when model responses drift from a reference snapshot or fail structural requirements.
    • Safety enforcement applies tagged test suites to ensure privacy and safety constraints are maintained across model updates.
    • Standardized reporting generates Markdown and JSON evaluation reports for CI/CD integration or human review.

    How it works

    1. Define scenarios by creating a JSON test suite containing inputs, expected tool calls, and text assertions.
    2. Capture run records from your application as JSONL files containing the actual model outputs and metadata.
    3. Execute the runner using the provided script to evaluate the captured records against your test suite.
    4. Analyze reports to identify failed assertions, missing tool calls, or unexpected schema changes.

    Frameworks & tools

    This skill is provider-neutral and works with any LLM stack that can export JSON/JSONL traces. It uses a Python-based execution script for local or CI/CD validation of structured data.

    Why this beats prompting it yourself

    Manual inspection of agent traces is slow and prone to human error. This skill replaces vague "vibes-based" testing with strict, repeatable assertions that block broken builds before they reach production.

    Use cases

    • Verify that a customer support agent calls the refund_policy tool when asked about returns.
    • Ensure a data extraction prompt always returns valid JSON matching a specific schema.
    • Confirm that a model update hasn't introduced regressions in a complex multi-step reasoning chain.

    Known limitations

    This runner is deterministic and does not perform semantic or LLM-based evaluation. It requires external capture of run records as it does not call model APIs directly.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    • Passed all security checks, Safe to install

    Listedtoday

    What's inside

    Frequently Asked Questions