New: UPI payments are live. Buyers in India can now pay for skills with UPI in INR -> Browse skills

    Browse The Skill Store

    14 skills found

    benchmarking ai agents beyond models

    by loreto

    Free

    Published AI benchmarks measure brains in jars. They test models in isolation or within a single reference harness — and then attribute all performance to the model. This skill teaches you to decompose agent performance into its two actual components: model capability and harness multiplier. The result is evaluations that predict real-world behavior instead of benchmark theater.

    2
    165.0(1)

    prompt engineer pro

    by Roy Yuen

    $8

    Professional prompt engineering, audit, and evaluation system for production-grade AI agents and workflows.

    3
    0

    production agent architect

    by Roy Yuen

    $6

    Architect, scaffold, and harden production-grade AI agents with battle-tested patterns and systematic evaluation.

    2
    2

    agent eval coverage audit

    by Roy Yuen

    $5

    Audit your AI agent's evaluation coverage to identify missing release gates and production risks.

    2
    0

    AI Eval & Test Suite Quality Gate

    by PubsProToolkit

    Free

    An adversarial gate that audits an AI eval or test suite — LLM-judge rubrics, datasets, regression tests, metrics — for gameable criteria, data leakage, missing edge cases, and non-determinism, then returns one PASS/REVISE/FAIL verdict.

    2
    1

    LLM Eval Framework Builder

    by StrategistKit

    $17

    Builds a complete LLM evaluation framework — quality dimensions, a golden dataset, code-based and model-graded rubric graders, judge calibration, and CI regression rules. Use when the user says build LLM evals, create a golden dataset, or set up LLM-as-judge. Do not use when they want to debug one bad model output, not build a repeatable measurement system.

    1
    0

    AI Feature Eval Writer

    by PubsProToolkit

    Free

    Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.

    1
    1

    rag eval

    by Ifásola

    $5

    Diagnose RAG bottlenecks with precision metrics (Recall, MRR, nDCG) to identify retrieval or ranking failures.

    2
    0

    Agent Harness Architect

    by PubsProToolkit

    Free

    Model quality is table stakes — the harness is where agents win or fail. This designs yours: it writes a structured, testable system prompt (role, tools, boundaries, method, output contract, failure handling) and maps every concern to the right layer — prompt, tool, guardrail, or evaluation — so the pieces reinforce each other instead of fighting.

    1
    0

    Skill Evaluation Harness Builder

    by Scar Tissue Systems

    $7.99

    Your skill works today. Will it work after the next model update? Build the harness that answers with numbers. Builds a standalone regression test harness with mechanical grading to verify skill behavior after model or code updates.

    1
    0

    🧠 Retrieval Trace Auditor

    by JustHandled Labs

    $15

    Audit real RAG evidence traces for missing expected sources, impossible citations, unsupported claims, stale evidence, unused context, and weak routing.

    1
    0

    Agent Evaluation & Regression Suite — Score Skills Before You Ship

    by Vivek K

    $15

    Designs a scored eval suite for an AI agent or SKILL.md skill, measures trigger accuracy, catches regressions against a baseline, and issues a PASS / PASS-WITH-WARNINGS / FAIL release gate.

    1
    0

    RAG Hallucination Root Cause Analyzer

    by heyhridyansh

    $14.99

    Diagnose RAG hallucinations, retrieval failures, and citation errors with a structured root-cause audit.

    1
    0

    Agent Observability Contract Designer

    by 香意國際

    $19.99

    Design a vendor-neutral telemetry contract for one AI agent workflow: traces, events, metrics, evaluations, redaction, dashboards, and alert hypotheses.

    1
    0