New: Skill bounties are live. Post a request, fund the bounty, and creators compete for 7 days to build it -> See open bounties

    Browse The Skill Store

    37 skills found

    synthesizing institutional knowledge

    by loreto

    $10

    Builds the organizational memory schema your AI agent needs to answer why — capturing decision provenance, causal chains, and event context that embedding-based retrieval permanently discards.

    10
    55.0(1)

    designing hybrid context layers

    by loreto

    $10

    Architects the right retrieval strategy for every query — teaching your agent when to use RAG, a knowledge graph, or a temporal index instead of defaulting to vector search for everything.

    6
    165.0(1)

    benchmarking ai agents beyond models

    by loreto

    Free

    Published AI benchmarks measure brains in jars. They test models in isolation or within a single reference harness — and then attribute all performance to the model. This skill teaches you to decompose agent performance into its two actual components: model capability and harness multiplier. The result is evaluations that predict real-world behavior instead of benchmark theater.

    2
    155.0(1)

    evaluating ai harness dimensions

    by loreto

    $10

    Evaluates AI coding agent platforms across five structural dimensions that determine real-world performance independently of model quality, so teams select on architectural fit rather than benchmark scores.

    3
    0

    prompt engineer pro

    by Roy Yuen

    $8

    Professional prompt engineering, audit, and evaluation system for production-grade AI agents and workflows.

    3
    0

    diagnosing rag failure modes

    by loreto

    $10

    RAG fails quietly. It retrieves documents, returns confident-looking answers, and misses the question entirely — because the question required connecting facts across documents, reasoning about sequence, or tracing causation. This skill gives you a five-question diagnostic checklist that classifies any failing query as either RAG-safe or structurally RAG-incompatible, then maps it to the specific failure pattern and the architectural fix that resolves it.

    5
    5

    production agent architect

    by Roy Yuen

    $6

    Architect, scaffold, and harden production-grade AI agents with battle-tested patterns and systematic evaluation.

    2
    2

    Skill Health Scanner

    by Markus Isaksson

    Free

    Instantly diagnose any skill or prompt and get a clear, prioritized report on what’s wrong and how to fix it — across any agent.

    2
    15

    agent eval coverage audit

    by Roy Yuen

    $5

    Audit your AI agent's evaluation coverage to identify missing release gates and production risks.

    2
    0

    Agensi Free Skill Explorer with Grok (v1.2.1)

    by Markus Isaksson

    Free

    Find and evaluate the best free Agensi marketplace skills for your specific development needs using Grok and MCP.

    2
    9

    steam api

    by Nicolas KLETHI

    $9

    Connect your agent to the Steam Web API to fetch player data, game libraries, and achievement statistics.

    2
    0

    AI Eval & Test Suite Quality Gate

    by PubsProToolkit

    Free

    An adversarial gate that audits an AI eval or test suite — LLM-judge rubrics, datasets, regression tests, metrics — for gameable criteria, data leakage, missing edge cases, and non-determinism, then returns one PASS/REVISE/FAIL verdict.

    2
    1

    Optimization Loop

    by Martin Gunderman

    $19

    Autonomous loop that iteratively modifies, evaluates, and selects the best version of any text resource — skills, prompts, or campaigns — using a modify-measure-keep/discard cycle.

    1
    1

    🧠 AI Memory Optimizer

    by Martin Gunderman

    $7

    Drastically reduce RAG costs and latency while improving retrieval accuracy through advanced memory architecture.

    2
    1

    LLM Eval Framework Builder

    by Arnstein Larsen

    $17

    You changed the prompt, tried four inputs, it looked better, you shipped — and three days later support tickets say outputs are worse for an entire class of inputs you didn't test

    1
    0

    Feature Intake & Triage

    by Indy Agent

    $5

    Evaluate any feature request with structured scoring and a Build / Investigate / Defer / Decline decision.

    2
    0

    normalize prompt set

    by Corey Jacobs

    Free

    Convert loose prompt sets into structured, target-ready records with variables, contracts, and eval cases.

    1
    2

    agent skill security auditor

    by Timoranjes

    $0

    Evaluate third-party agent skills for command injection, prompt injection, and data exfiltration before installation.

    2
    0

    rag failure diagnostics

    by Kaymue

    Free

    Diagnose broken RAG systems. 8 failure categories: chunking, embeddings, retrieval, reranking, hallucination. Recall@k measurement.

    2
    2

    AI Feature Eval Writer

    by PubsProToolkit

    Free

    Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.

    1
    1

    RAG Failure Diagnostics & Architect

    by Arnstein Larsen

    $7

    A retrieval architect that diagnoses why RAG returns confident-but-wrong answers, picks the right context architecture (RAG vs knowledge graph vs structured/temporal retrieval) instead of defaulting to vector search, and designs the institutional-memory schema embeddings throw away.

    1
    0

    rag eval

    by Ifásola

    $5

    Diagnose RAG bottlenecks with precision metrics (Recall, MRR, nDCG) to identify retrieval or ranking failures.

    2
    0

    Agent Eval Harness

    by Echo Rose

    $5

    Agent Eval Harness - A Premium AI Agent Skill

    1
    0

    agent skill regression tester

    by Timoranjes

    Free

    Teaches AI coding agents to build and run automated regression tests for SKILL.md files. When you update a skill that your team depends on, you need to know it still works — not just that it "looks ri

    2
    3