Ship better AI in 30 seconds. Browse 2,000+ expert-built and security scanned skills -> Browse skills

    Browse The Skill Store

    4 skills found

    benchmarking ai agents beyond models

    by loreto

    Free

    Published AI benchmarks measure brains in jars. They test models in isolation or within a single reference harness — and then attribute all performance to the model. This skill teaches you to decompose agent performance into its two actual components: model capability and harness multiplier. The result is evaluations that predict real-world behavior instead of benchmark theater.

    1
    155.0(1)

    evaluating ai harness dimensions

    by loreto

    $10

    Evaluates AI coding agent platforms across five structural dimensions that determine real-world performance independently of model quality, so teams select on architectural fit rather than benchmark scores.

    3
    0No reviews

    Kubernetes Config Error Detective

    by Shandra

    $50

    Audits Kubernetes manifests, Helm values, deployment logs, and service configs to detect configuration errors and produce safe, reviewable fix plans.

    2
    0No reviews

    Fleet Scale Migration Orchestrator Human Supervised

    by John Barros

    $39

    Orchestrate human-supervised code migrations across repository fleets with verifier loops and judge review gates.

    2
    0No reviews