More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    RAG Retrieval Release Gate

    1

    Decide whether a RAG retrieval change is ready to ship using an evidence-labelled case ledger, hard negatives, acceptance checks, and a limited-release verdict.

    $12

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    RAG Retrieval Release Gate

    RAG Retrieval Release Gate

    Example session with this skill installed

    We are considering a new reranker for our internal support RAG. It will run after hybrid search. We need to know whether it is safe to release to 10% of users.

    Corpus rules

    • Every policy article has tenant_id, effective_date, and status fields.
    • Current policies have status=active; retired policies have status=retired.
    • The tenant filter must always remain in place.

    Case A
    Query: How long can an EU customer keep an export link active?
    Expected source: EU-EXPORT-2025, active policy for tenant northwind.
    Current results: EU-EXPORT-2025 rank 1; EU-EXPORT-2023 retired rank 4; US-EXPORT-2025 rank 7.
    Candidate results: EU-EXPORT-2023 retired rank 1; EU-EXPORT-2025 rank 5; US-EXPORT-2025 rank 8.

    Case B
    Query: Can a customer remove an audit log entry?
    Expected source: AUDIT-IMMUTABILITY-01, active for tenant northwind.
    Candidate results: AUDIT-IMMUTABILITY-01 rank 2; LOG-RETENTION-04 rank 3; AUDIT-IMMUTABILITY-01 for tenant contoso rank 5.

    The candidate answer text is not available. Do not query our systems or invent metrics. Produce a retrieval release gate with required checks, hard negatives, and a decision.

    • Read your context and instructions
    • Compiled the rag retrieval release
    • Generated the document

    RELEASE DECISION

    BLOCK. The candidate promotes a retired policy to rank 1 for Case A while the required active policy falls to rank 5. It also exposes a cross-tenant near-match in Case B. The supplied evidence does not support a 10% rollout.

    CHANGE BOUNDARY

    Confirmed: The proposed change is a new reranker after hybrid search. The requested rollout is 10%.
    Confirmed: Tenant filtering and active-policy selection are mandatory.
    Unknown: Candidate answer text, filter trace, score breakdown, latency, and behavior outside these two cases were not supplied.

    EVALUATION CASE LEDGER

    Case A — EU export-link duration
    Expected: EU-EXPORT-2025, tenant northwind, active.
    Observed: Retired EU-EXPORT-2023 is rank 1; the expected source is rank 5.
    Status: BLOCKER. The wrong policy version outranks the required evidence.

    Case B — audit-log deletion
    Expected: AUDIT-IMMUTABILITY-01, tenant northwind, active.
    Observed: The expected source is rank 2; the same document ID from tenant contoso appears at rank 5.
    Status: CONDITIONAL. The correct source is present, but tenant isolation is not proven.

    HARD NEGATIVES

    EU-EXPORT-2023 looks relevant but is retired.
    US-EXPORT-2025 is the wrong region for Case A.
    AUDIT-IMMUTABILITY-01 from contoso is forbidden for a northwind query.
    The candidate must prove tenant_id=northwind and status=active are applied before reranking.

    FINDINGS

    Confirmed: Case A fails the expected-source requirement.
    Supported: The reranker may be over-weighting lexical similarity or stale content. Score and filter traces are required to confirm this.
    Unknown: Recall, precision, latency, and answer faithfulness cannot be calculated from two result sets.

    ACCEPTANCE CHECKS

    1. EU-EXPORT-2025 must outrank every retired or non-EU policy for Case A.
    2. A retired result cannot be eligible as answer evidence when an active version exists.
    3. A northwind query cannot retrieve content from another tenant.
    4. Add at least six labelled cases for versioning, tenant boundaries, numeric limits, and no-answer requests.
    5. Record the query, applied filters, result IDs, ranks, and release version for every case.

    REGRESSION AND RELEASE PLAN

    Keep both supplied cases in the regression set. After every acceptance check passes, run a limited rollout with a named human approver, privacy-safe retrieval logging, and an immediate rollback trigger for retired or cross-tenant evidence. This review does not authorize deployment.

    rag-retrieval-release-gate.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Decide whether a reranker or hybrid-search change is safe to release.Catch retired, wrong-region, or cross-tenant sources before rollout.Turn retrieval traces into acceptance checks and a regression set.Define a limited rollout and rollback plan for a RAG retrieval change.

    About this skill

    Decide whether a retrieval change is ready to ship RAG Retrieval Release Gate is for the risky moment before you release a change to chunking, embeddings, metadata filters, hybrid search, reranking, or an index. A few promising answers are not enough. A retrieval change can silently promote a retired policy, cross a tenant boundary, or retrieve a near-match that supports the wrong conclusion.

    What you provide Supply the proposed change and the evidence you have: real user queries, expected document IDs or source classes, current and candidate top-k results, excerpts, ranks, filters, corpus notes, and release constraints.

    What you receive An evidence-labelled evaluation case ledger. An expected-source map and a list of hard negatives. Acceptance checks and a reusable regression set. Rollout, monitoring, rollback, and approval requirements. One verdict: BLOCK, CONDITIONAL, READY FOR LIMITED RELEASE, or INSUFFICIENT EVIDENCE. Built for concrete retrieval changes Use it to review a reranker, hybrid-search change, chunking update, embedding migration, metadata filter, or index refresh before users depend on it.

    Safe by design The skill works only from evidence supplied in the conversation. It does not query your vector database, change an index, invent recall numbers, or authorize a live release.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 20 days ago

    • Passed all security checks, Safe to install

    Listed20 days ago

    What's inside

    Frequently Asked Questions