More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    qikly

    2

    Enforce spec-driven Python testing by withholding acceptance criteria from the coding agent to ensure tests can fail.

    Free

    2 installs5.0 (2 reviews)
    qikly

    qikly

    Example session with this skill installed

    I have a module pricing.py and a test suite in test_pricing.py that passes. I want to know whether those tests would actually catch a bug, or whether they just agree with the code.

    • Read your context and instructions
    • Compiled the qikly

    I will plant one fault at a time in a copy of your code and run your existing
    tests against each one. This is local: no model call, no network, and nothing
    of yours is modified.

    Your suite first: 7 passed in 0.06s.

    Suite score for pricing.py

    4 of 8 planted faults caught (50%).

    Read that as a floor rather than a verdict. It says how much of the code that
    is here your tests would notice being changed. It cannot say anything about
    a rule nobody implemented, because there is nothing there to break, so a high
    score is not a clean bill of health.

    What it missed

    Change made to your codeLineKind
    >= became >11arithmetic
    <= became <25arithmetic
    10 became 114arithmetic
    guard disabled25validation

    What it caught

    1 to 2 at line 12, 1 to 2 and 2 to 3 at line 18, 0 to 1 at line 25.

    What the misses mean

    Three of the four are boundaries. The suite exercises quantities of 2 and 20
    but never 10, which is the bulk-discount threshold itself, so moving the
    threshold or flipping >= to > changes nothing it checks. That is the gap a
    suite written from the code cannot see: it was written by reading an
    implementation that had already decided where the boundary goes.

    Next steps

    1. Add a case at exactly quantity == 10.
    2. Decide, in writing, whether 10 is included, then test both sides of it.
    3. Run qikly --scaffold pricing.py to write that decision down as a
      specification the test writer sees and the coding agent does not.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    Developers often struggle with test suites that pass consistently but fail to catch real bugs. This happens when tests are written by reading the implementation, essentially codifying existing behavior rather than verifying a contract. This skill implements a specification-driven testing workflow for Python modules, ensuring tests are honest and capable of failing.

    What it does

    • Specification splitting divides requirements from acceptance criteria so the coding agent never sees the "answer key."
    • Automated scaffolding reads Python function signatures to generate task files and test structures.
    • Mutation scoring plants faults in existing code to measure how many a suite actually catches.
    • Validation checks identify unreachable criteria, ambiguous requirements, and missing data fixtures before you spend tokens.
    • Refinement loops iterate on code patches until the implementation satisfies the withheld criteria.

    How it works

    1. Scaffold the project using qikly --scaffold to generate a task file from your Python signatures.
    2. Define requirements (decisions) and acceptance criteria (consequences) in the task file, providing sample data inputs.
    3. Validate the setup for free using --validate to catch logic errors or unreachable code paths.
    4. Execute the testing run where the skill coordinates a coding agent and a testing agent in isolation.

    Frameworks & tools

    This tool is built specifically for Python environments and integrates with pytest. It supports models from Anthropic, OpenAI, and Google Gemini via API keys.

    Why this beats prompting it yourself

    Writing tests via standard prompting usually results in "consistent" code where the agent justifies its own bugs. This skill enforces a structural wall between the spec and the implementation that simple prompting cannot maintain, ensuring tests actually guard against regressions.

    Use cases

    • Auditing a legacy test suite to see if it actually catches logical faults.
    • Writing test-driven code from a spec without letting the coder see the edge-case answers.
    • Resolving ambiguities in financial or mathematical rounding rules before they reach production.
    • Generating comprehensive test fixtures for complex boundary conditions.

    Known limitations

    Supports Python modules only. Requires qikly version 0.5.4 or later. Convergence varies by model and is not a guarantee of correctness: a run can finish without converging, and the report says so when it does.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    5.0
    2 reviews
    5
    2
    4
    0
    3
    0
    2
    0
    1
    0

    2 people have installed this skill.

    Trust & safety

    Security scanned

    Verified clean 1 day ago

    • Free to download with an account

    Listed1 day ago

    What's inside

    Frequently Asked Questions