More screenshots
Works with the AI tools you already use
qikly
Enforce spec-driven Python testing by withholding acceptance criteria from the coding agent to ensure tests can fail.
Free
qikly
Example session with this skill installed
I have a module pricing.py and a test suite in test_pricing.py that passes. I want to know whether those tests would actually catch a bug, or whether they just agree with the code.
- Read your context and instructions
- Compiled the qikly
I will plant one fault at a time in a copy of your code and run your existing
tests against each one. This is local: no model call, no network, and nothing
of yours is modified.
Your suite first: 7 passed in 0.06s.
Suite score for pricing.py
4 of 8 planted faults caught (50%).
Read that as a floor rather than a verdict. It says how much of the code that
is here your tests would notice being changed. It cannot say anything about
a rule nobody implemented, because there is nothing there to break, so a high
score is not a clean bill of health.
What it missed
| Change made to your code | Line | Kind |
|---|---|---|
>= became > | 11 | arithmetic |
<= became < | 25 | arithmetic |
10 became 11 | 4 | arithmetic |
| guard disabled | 25 | validation |
What it caught
1 to 2 at line 12, 1 to 2 and 2 to 3 at line 18, 0 to 1 at line 25.
What the misses mean
Three of the four are boundaries. The suite exercises quantities of 2 and 20
but never 10, which is the bulk-discount threshold itself, so moving the
threshold or flipping >= to > changes nothing it checks. That is the gap a
suite written from the code cannot see: it was written by reading an
implementation that had already decided where the boundary goes.
Next steps
- Add a case at exactly
quantity == 10. - Decide, in writing, whether 10 is included, then test both sides of it.
- Run
qikly --scaffold pricing.pyto write that decision down as a
specification the test writer sees and the coding agent does not.
Connects securely to your tools. The creator never sees your data.
About this skill
Developers often struggle with test suites that pass consistently but fail to catch real bugs. This happens when tests are written by reading the implementation, essentially codifying existing behavior rather than verifying a contract. This skill implements a specification-driven testing workflow for Python modules, ensuring tests are honest and capable of failing.
What it does
- Specification splitting divides requirements from acceptance criteria so the coding agent never sees the "answer key."
- Automated scaffolding reads Python function signatures to generate task files and test structures.
- Mutation scoring plants faults in existing code to measure how many a suite actually catches.
- Validation checks identify unreachable criteria, ambiguous requirements, and missing data fixtures before you spend tokens.
- Refinement loops iterate on code patches until the implementation satisfies the withheld criteria.
How it works
- Scaffold the project using
qikly --scaffoldto generate a task file from your Python signatures. - Define requirements (decisions) and acceptance criteria (consequences) in the task file, providing sample data inputs.
- Validate the setup for free using
--validateto catch logic errors or unreachable code paths. - Execute the testing run where the skill coordinates a coding agent and a testing agent in isolation.
Frameworks & tools
This tool is built specifically for Python environments and integrates with pytest. It supports models from Anthropic, OpenAI, and Google Gemini via API keys.
Why this beats prompting it yourself
Writing tests via standard prompting usually results in "consistent" code where the agent justifies its own bugs. This skill enforces a structural wall between the spec and the implementation that simple prompting cannot maintain, ensuring tests actually guard against regressions.
Use cases
- Auditing a legacy test suite to see if it actually catches logical faults.
- Writing test-driven code from a spec without letting the coder see the edge-case answers.
- Resolving ambiguities in financial or mathematical rounding rules before they reach production.
- Generating comprehensive test fixtures for complex boundary conditions.
Known limitations
Supports Python modules only. Requires qikly version 0.5.4 or later. Convergence varies by model and is not a guarantee of correctness: a run can finish without converging, and the report says so when it does.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
2 people have installed this skill.
Trust & safety
Security scanned
Verified clean 1 day ago
- Free to download with an account