Get expert-level AI output in 30 seconds. Browse 2,000+ expert-built and security scanned skills -> Browse skills

    Browse The Skill Store

    3 skills found

    benchmarking ai agents beyond models

    by loreto

    Free

    Published AI benchmarks measure brains in jars. They test models in isolation or within a single reference harness — and then attribute all performance to the model. This skill teaches you to decompose agent performance into its two actual components: model capability and harness multiplier. The result is evaluations that predict real-world behavior instead of benchmark theater.

    2
    155.0(1)

    evaluating ai harness dimensions

    by loreto

    $10

    Evaluates AI coding agent platforms across five structural dimensions that determine real-world performance independently of model quality, so teams select on architectural fit rather than benchmark scores.

    3
    0No reviews

    Eval Harness Builder

    by Scar Tissue Systems

    $10

    Your skill works today. Will it work after the next model update? Build the harness that answers with numbers. Builds a standalone regression test harness with mechanical grading to verify skill behavior after model or code updates.

    1
    0No reviews