New: Skill bounties are live. Post a request, fund the bounty, and creators compete for 7 days to build it -> See open bounties

    Browse The Skill Store

    63 skills found

    LLM Eval Framework Builder

    by Arnstein Larsen

    $17

    You changed the prompt, tried four inputs, it looked better, you shipped — and three days later support tickets say outputs are worse for an entire class of inputs you didn't test

    1
    0

    Feature Intake & Triage

    by Indy Agent

    $5

    Evaluate any feature request with structured scoring and a Build / Investigate / Defer / Decline decision.

    2
    0

    Prompt Dataset Builder

    by TopAgent

    Free

    Build and curate high-quality prompt datasets for fine-tuning and eval — deduped and labeled.

    1
    58

    🤖 3. AI Readiness Score

    by Martin Gunderman

    $7

    Evaluate company AI maturity across 6 dimensions with weighted scoring, radar charts, and a GDPR risk audit.

    1
    0

    agent skill security auditor

    by Timoranjes

    $0

    Evaluate third-party agent skills for command injection, prompt injection, and data exfiltration before installation.

    2
    0

    normalize prompt set

    by GTDataworks

    Free

    Convert loose prompt sets into structured, target-ready records with variables, contracts, and eval cases.

    1
    2

    onkosten categorisatie

    by Nex AI

    $7

    Wijst onkostenregels uit een CSV-export automatisch toe aan Belgische MAR-voorbeeldrekeningen op basis van trefwoorden, met review-flags voor twijfelgevallen.

    4
    0

    RAG Failure Diagnostics & Architect

    by Arnstein Larsen

    $7

    A retrieval architect that diagnoses why RAG returns confident-but-wrong answers, picks the right context architecture (RAG vs knowledge graph vs structured/temporal retrieval) instead of defaulting to vector search, and designs the institutional-memory schema embeddings throw away.

    1
    0

    rag failure diagnostics

    by Kaymue

    Free

    Diagnose broken RAG systems. 8 failure categories: chunking, embeddings, retrieval, reranking, hallucination. Recall@k measurement.

    2
    2

    AI Feature Eval Writer

    by PubsProToolkit

    Free

    Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.

    1
    1

    rag eval

    by Ifásola

    $5

    Diagnose RAG bottlenecks with precision metrics (Recall, MRR, nDCG) to identify retrieval or ranking failures.

    2
    0

    Agent Eval Harness

    by Echo Rose

    $5

    Agent Eval Harness - A Premium AI Agent Skill

    1
    0

    agent skill regression tester

    by Timoranjes

    Free

    Teaches AI coding agents to build and run automated regression tests for SKILL.md files. When you update a skill that your team depends on, you need to know it still works — not just that it "looks ri

    2
    3

    Opencv5 Native Dnn Inference Migration Scaffold

    by John Barros

    $59

    Evaluate and plan the migration of vision inference pipelines to native OpenCV 5 DNN CPU execution.

    2
    0

    Agent Harness Architect

    by PubsProToolkit

    Free

    Model quality is table stakes — the harness is where agents win or fail. This designs yours: it writes a structured, testable system prompt (role, tools, boundaries, method, output contract, failure handling) and maps every concern to the right layer — prompt, tool, guardrail, or evaluation — so the pieces reinforce each other instead of fighting.

    1
    0

    Agent Memory Design Planner

    by PromptWagon

    $10

    Designs practical memory architectures for AI assistants, agents, copilots, automations, and workflows, including memory schemas, retention rules, update policies, retrieval keys, summary formats, privacy boundaries, conflict handling, user preference memory, project memory, task memory, and audit-ready memory governance notes for builders.

    1
    0

    RAG Knowledge Base Auditor

    by PromptWagon

    $10

    Reviews document sets, source quality, chunking logic, metadata, retrieval coverage, citation traceability, answer grounding, source gaps, stale content, duplicate content, and failure patterns for RAG knowledge-base chatbots. Helps AI, product, support, governance, and engineering teams diagnose common and costly RAG quality problems before deployment or after incidents.

    1
    0

    Model Evaluation Report Builder

    by PromptWagon

    $10

    Turns model test results, prompts, outputs, benchmarks, scoring notes, evaluation datasets, failure examples, comparison results, and reviewer observations into clear model evaluation reports with findings, recommendations, evidence gaps, deployment considerations, and repeatable evaluation documentation for AI teams.

    1
    0

    Agentic Engineering Patterns Library

    by PubsProToolkit

    Free

    A deep reference library of production agent patterns — orchestration, context, tool design, failure and recovery, oversight, and evaluation. Every pattern states when it applies, when it's the wrong answer, what it costs, and the failure it prevents. Seven reference files, not a checklist.

    1
    0

    Skill Evaluation Harness Builder

    by Scar Tissue Systems

    $7.99

    Your skill works today. Will it work after the next model update? Build the harness that answers with numbers. Builds a standalone regression test harness with mechanical grading to verify skill behavior after model or code updates.

    1
    0

    🔁 Automation Acceptance Replay Harness

    by JustHandled Labs

    $19

    Replay automation traces against explicit rules for required steps, order, retries, approvals, idempotency, external side effects, duration, and terminal success.

    1
    0

    Agent Security and Red Team Pattern Library

    by PubsProToolkit

    Free

    A defensive catalog of ~39 agent security attack patterns across every surface — injection, tool abuse, exfiltration, memory poisoning, multi-agent trust, retrieval poisoning — each with a detection signal, a concrete defense, and a severity. Nine reference files plus a threat-model worksheet. For hardening agents you own.

    2
    0

    🧠 Retrieval Trace Auditor

    by JustHandled Labs

    $15

    Audit real RAG evidence traces for missing expected sources, impossible citations, unsupported claims, stale evidence, unused context, and weak routing.

    1
    0

    Checkout & Payment UI Optimization Agent

    by Shandra

    $7.99

    Audits and optimizes checkout and payment experiences for e-commerce stores, SaaS subscriptions, marketplaces, booking platforms, and digital products. Evaluates cart review, pricing transparency, guest checkout, shipping, billing, payment methods, discounts, trust, validation, declines, uncertain payment states, order confirmation, accessibility, analytics, abandoned-checkout recovery, and implementation readiness.

    1
    0