Ship better AI in 30 seconds. Browse 2,000+ expert-built and security scanned skills -> Browse skills

    Browse The Skill Store

    17 skills found

    designing hybrid context layers

    by loreto

    $10

    Architects the right retrieval strategy for every query — teaching your agent when to use RAG, a knowledge graph, or a temporal index instead of defaulting to vector search for everything.

    5
    165.0(1)

    benchmarking ai agents beyond models

    by loreto

    Free

    Published AI benchmarks measure brains in jars. They test models in isolation or within a single reference harness — and then attribute all performance to the model. This skill teaches you to decompose agent performance into its two actual components: model capability and harness multiplier. The result is evaluations that predict real-world behavior instead of benchmark theater.

    1
    155.0(1)

    prompt engineer pro

    by Roy Yuen

    $8

    Professional prompt engineering, audit, and evaluation system for production-grade AI agents and workflows.

    3
    0No reviews

    diagnosing rag failure modes

    by loreto

    $10

    RAG fails quietly. It retrieves documents, returns confident-looking answers, and misses the question entirely — because the question required connecting facts across documents, reasoning about sequence, or tracing causation. This skill gives you a five-question diagnostic checklist that classifies any failing query as either RAG-safe or structurally RAG-incompatible, then maps it to the specific failure pattern and the architectural fix that resolves it.

    4
    5No reviews

    production agent architect

    by Roy Yuen

    $6

    Architect, scaffold, and harden production-grade AI agents with battle-tested patterns and systematic evaluation.

    2
    2No reviews

    agent eval coverage audit

    by Roy Yuen

    $5

    Audit your AI agent's evaluation coverage to identify missing release gates and production risks.

    2
    0No reviews

    AI Eval & Test Suite Quality Gate

    by PubsProToolkit

    $14

    An adversarial gate that audits an AI eval or test suite — LLM-judge rubrics, datasets, regression tests, metrics — for gameable criteria, data leakage, missing edge cases, and non-determinism, then returns one PASS/REVISE/FAIL verdict.

    2
    1No reviews

    🧠 AI Memory Optimizer

    by Martin Gunderman

    $7

    Drastically reduce RAG costs and latency while improving retrieval accuracy through advanced memory architecture.

    2
    1No reviews

    LLM Eval Framework Builder

    by Arnstein Larsen

    $17

    You changed the prompt, tried four inputs, it looked better, you shipped — and three days later support tickets say outputs are worse for an entire class of inputs you didn't test

    1
    0No reviews

    normalize prompt set

    by Corey Jacobs

    Free

    Convert loose prompt sets into structured, target-ready records with variables, contracts, and eval cases.

    1
    2No reviews

    RAG Failure Diagnostics & Architect

    by Arnstein Larsen

    $7

    A retrieval architect that diagnoses why RAG returns confident-but-wrong answers, picks the right context architecture (RAG vs knowledge graph vs structured/temporal retrieval) instead of defaulting to vector search, and designs the institutional-memory schema embeddings throw away.

    1
    0No reviews

    rag failure diagnostics

    by Kaymue

    Free

    Diagnose broken RAG systems. 8 failure categories: chunking, embeddings, retrieval, reranking, hallucination. Recall@k measurement.

    2
    2No reviews

    rag eval

    by Ifásola

    $5

    Diagnose RAG bottlenecks with precision metrics (Recall, MRR, nDCG) to identify retrieval or ranking failures.

    2
    0No reviews

    AI Feature Eval Writer

    by PubsProToolkit

    $14

    Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.

    1
    0No reviews

    Agent Eval Harness

    by Echo Rose

    $5

    Agent Eval Harness - A Premium AI Agent Skill

    1
    0No reviews

    Agent Harness Architect

    by PubsProToolkit

    $14

    Model quality is table stakes — the harness is where agents win or fail. This designs yours: it writes a structured, testable system prompt (role, tools, boundaries, method, output contract, failure handling) and maps every concern to the right layer — prompt, tool, guardrail, or evaluation — so the pieces reinforce each other instead of fighting.

    1
    0No reviews

    Eval Harness Builder

    by Scar Tissue Systems

    $12

    Your skill works today. Will it work after the next model update? Build the harness that answers with numbers. Builds a standalone regression test harness with mechanical grading to verify skill behavior after model or code updates.

    1
    0No reviews