Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    RAG Eval Suite

    1

    Production-grade RAG evaluation suite with 19 metrics, safety gates, and noise-aware regression testing.

    $15

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    RAG Eval Suite

    RAG Eval Suite

    Example session with this skill installed

    I need to set up a production-grade eval suite for my LangChain RAG app. I have a retriever and a generator function. How do I handle regression testing without getting false positives from noise?

    • Read your context and instructions
    • Compiled the rag eval suite

    I have mapped out a production-grade evaluation suite for your LangChain application. The implementation focuses on isolated component testing, a wired RAG Triad, and a noise-aware regression harness to ensure that only statistically significant quality drops flag a failure.

    RAG Evaluation Suite Implementation
    Layer Metrics Scripts
    Component Contextual Recall, Contextual Precision (retriever) · Faithfulness, Answer Relevancy (generator) eval_retriever.py, eval_generator.py
    Pipeline RAG Triad — Contextual Relevancy, Faithfulness, Answer Relevancy eval_pipeline.py
    Application Correctness, Completeness, Style (G-Eval) eval_application.py
    Safety Scope Adherence, System Prompt Leakage, PII Leakage, Toxicity eval_safety.py
    Operations P95 Latency, TTFT, Cost per Query, Reliability (success/error rate) eval_operations.py
    Regression Noise-floor adjusted delta monitoring noise_baseline.py, compare.py
    Noise-Aware Regression Strategy

    To prevent false positives, we will not perform naive diffing. Instead:

    Run noise_baseline.py on your current stable version 10 times to establish the standard deviation for each metric.
    compare.py sets the regression threshold at 2 standard deviations above that noise band, per metric.
    Only deltas exceeding this noise band will trigger a regression alert in your CI/CD pipeline — everything inside the band is reported as noise, not a failure.
    yaml

    config.yaml — point this at your LangChain retriever/generator/pipeline functions

    judge_model: gpt-4o-mini
    functions:
    retriever_fn: my_app.retriever:retrieve
    generator_fn: my_app.generator:generate
    pipeline_fn: my_app.pipeline:run
    golden_datasets:
    retriever: examples/retriever_golden.json
    generator: examples/generator_golden.json
    pipeline: examples/pipeline_golden.json
    bash

    Establish the noise band once (or whenever the judge model changes)

    python scripts/noise_baseline.py --config config.yaml --runs 10 --out results/noise_band.json

    Run the full offline suite to get your baseline

    python scripts/run_suite.py --config config.yaml --out results/baseline.json

    After a change (chunking, prompt, model, reranker), rerun and compare

    python scripts/run_suite.py --config config.yaml --out results/candidate.json
    python scripts/compare.py --baseline results/baseline.json
    --candidate results/candidate.json --noise-band results/noise_band.json

    Next steps

    Fill in config.yaml with dotted paths to your LangChain retriever, generator, and pipeline functions, and point it at your golden dataset files (see config.example.yaml and examples/sample_golden_dataset.json for the expected shape).
    Run python scripts/noise_baseline.py --config config.yaml once to establish your judge model's variance floor — this is what makes compare.py's regression calls trustworthy instead of noisy.
    Run python scripts/run_suite.py --config config.yaml to get your baseline, then wire compare.py into CI so it runs on every PR that touches retrieval, prompts, or the model.

    config.yaml

    Connects securely to your tools. The creator never sees your data.

    What you get

    Automate RAG regression testing in CI/CD pipelines.Identify if retrieval or generation is causing pipeline failures.Detect production quality drift using LangSmith integration.Enforce safety gates for PII and toxicity before deployment.

    About this skill

    The problem

    Most RAG evaluation is shallow, often relying on a single faithfulness score run once. This fails to catch regressions in retrieval, misses production drift, and ignores the inherent noise in LLM-as-a-judge scoring.

    What it does

    • Executes a layered evaluation suite covering isolated components (retriever/generator), wired pipelines (RAG Triad), and application-level quality (G-Eval).
    • Implements safety gates for toxicity, PII leakage, and system prompt protection to block unsafe deployments.
    • Measures operational health including P95 latency, TTFT, cost per query, and reliability rates.
    • Provides a noise-aware regression harness that distinguishes real quality drops from standard LLM variance.
    • Integrates online production monitoring via LangSmith to detect distribution drift against offline baselines.

    Frameworks & tools

    Python, DeepEval, LangSmith, and CLI-based automation. Compatible with any RAG stack including LangChain, LlamaIndex, or custom implementations.

    Why this beats prompting it yourself

    Writing a prompt for 'faithfulness' is easy; building a production-grade test harness that handles claim decomposition, noise-floor calculation, and multi-stage evaluation is complex. This skill provides the architectural spine and statistical rigor needed to trust your RAG metrics.

    Use cases

    • Setting up a CI/CD gate that blocks RAG deployments if safety or quality metrics drop.
    • Benchmarking new embedding models or rerankers against a human-verified golden dataset.
    • Monitoring live production traffic for quality drift without requiring golden answers.
    • Decomposing RAG failures to identify if the retriever or the generator is the bottleneck.

    Known limitations

    Requires a clear seam between retriever and generator functions for component-level metrics. LLM-judge scripts require an OpenAI or similar API key for scoring.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 17 days ago

    • Passed all security checks, Safe to install

    Needs access to

    OpenAI
    Anthropic
    Langchain

    Listed17 days ago

    What's inside

    Frequently Asked Questions