New: Skill bounties are live. Post a request, fund the bounty, and creators compete for 7 days to build it -> See open bounties

    Browse The Skill Store

    5 skills found

    AI Eval & Test Suite Quality Gate

    by PubsProToolkit

    Free

    An adversarial gate that audits an AI eval or test suite — LLM-judge rubrics, datasets, regression tests, metrics — for gameable criteria, data leakage, missing edge cases, and non-determinism, then returns one PASS/REVISE/FAIL verdict.

    2
    1

    AI Feature Eval Writer

    by PubsProToolkit

    Free

    Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.

    1
    1

    Model Evaluation Report Builder

    by PromptWagon

    $10

    Turns model test results, prompts, outputs, benchmarks, scoring notes, evaluation datasets, failure examples, comparison results, and reviewer observations into clear model evaluation reports with findings, recommendations, evidence gaps, deployment considerations, and repeatable evaluation documentation for AI teams.

    1
    0

    Prompt Dataset Builder

    by TopAgent

    Free

    Build and curate high-quality prompt datasets for fine-tuning and eval — deduped and labeled.

    1
    0

    Computer Vision Dataset Quality Auditor

    by Shandra

    $9.99

    Audits image and video datasets for exact and near duplicates, cross-split leakage, class imbalance, suspected labeling errors, annotation inconsistencies, blur, low resolution, corruption, compression artifacts, visual shortcuts, source bias, insufficient diversity, synthetic-data weaknesses, and deployment-domain gaps. Produces an evidence-based quality report with class-balance analysis, labeling risks, severity-ranked findings, and a prioritized remediation and re-audit plan.

    1
    0