New: UPI payments are live. Buyers in India can now pay for skills with UPI in INR -> Browse skills

    Browse The Skill Store

    4 skills found

    AI Eval & Test Suite Quality Gate

    by PubsProToolkit

    Free

    An adversarial gate that audits an AI eval or test suite — LLM-judge rubrics, datasets, regression tests, metrics — for gameable criteria, data leakage, missing edge cases, and non-determinism, then returns one PASS/REVISE/FAIL verdict.

    2
    1

    LLM Eval Framework Builder

    by StrategistKit

    $17

    Builds a complete LLM evaluation framework — quality dimensions, a golden dataset, code-based and model-graded rubric graders, judge calibration, and CI regression rules. Use when the user says build LLM evals, create a golden dataset, or set up LLM-as-judge. Do not use when they want to debug one bad model output, not build a repeatable measurement system.

    1
    0

    AI Feature Eval Writer

    by PubsProToolkit

    Free

    Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.

    1
    1

    Skill Evaluation Harness Builder

    by Scar Tissue Systems

    $7.99

    Your skill works today. Will it work after the next model update? Build the harness that answers with numbers. Builds a standalone regression test harness with mechanical grading to verify skill behavior after model or code updates.

    1
    0