Works with the AI tools you already use
RAG Eval Suite
Production-grade RAG evaluation suite with 19 metrics, safety gates, and noise-aware regression testing.
$15
RAG Eval Suite
Example session with this skill installed
I need to set up a production-grade eval suite for my LangChain RAG app. I have a retriever and a generator function. How do I handle regression testing without getting false positives from noise?
- Read your context and instructions
- Compiled the rag eval suite
I have mapped out a production-grade evaluation suite for your LangChain application. The implementation focuses on isolated component testing, a wired RAG Triad, and a noise-aware regression harness to ensure that only statistically significant quality drops flag a failure.
RAG Evaluation Suite Implementation
Layer Metrics Scripts
Component Contextual Recall, Contextual Precision (retriever) · Faithfulness, Answer Relevancy (generator) eval_retriever.py, eval_generator.py
Pipeline RAG Triad — Contextual Relevancy, Faithfulness, Answer Relevancy eval_pipeline.py
Application Correctness, Completeness, Style (G-Eval) eval_application.py
Safety Scope Adherence, System Prompt Leakage, PII Leakage, Toxicity eval_safety.py
Operations P95 Latency, TTFT, Cost per Query, Reliability (success/error rate) eval_operations.py
Regression Noise-floor adjusted delta monitoring noise_baseline.py, compare.py
Noise-Aware Regression Strategy
To prevent false positives, we will not perform naive diffing. Instead:
Run noise_baseline.py on your current stable version 10 times to establish the standard deviation for each metric.
compare.py sets the regression threshold at 2 standard deviations above that noise band, per metric.
Only deltas exceeding this noise band will trigger a regression alert in your CI/CD pipeline — everything inside the band is reported as noise, not a failure.
yaml
config.yaml — point this at your LangChain retriever/generator/pipeline functions
judge_model: gpt-4o-mini
functions:
retriever_fn: my_app.retriever:retrieve
generator_fn: my_app.generator:generate
pipeline_fn: my_app.pipeline:run
golden_datasets:
retriever: examples/retriever_golden.json
generator: examples/generator_golden.json
pipeline: examples/pipeline_golden.json
bash
Establish the noise band once (or whenever the judge model changes)
python scripts/noise_baseline.py --config config.yaml --runs 10 --out results/noise_band.json
Run the full offline suite to get your baseline
python scripts/run_suite.py --config config.yaml --out results/baseline.json
After a change (chunking, prompt, model, reranker), rerun and compare
python scripts/run_suite.py --config config.yaml --out results/candidate.json
python scripts/compare.py --baseline results/baseline.json
--candidate results/candidate.json --noise-band results/noise_band.json
Next steps
Fill in config.yaml with dotted paths to your LangChain retriever, generator, and pipeline functions, and point it at your golden dataset files (see config.example.yaml and examples/sample_golden_dataset.json for the expected shape).
Run python scripts/noise_baseline.py --config config.yaml once to establish your judge model's variance floor — this is what makes compare.py's regression calls trustworthy instead of noisy.
Run python scripts/run_suite.py --config config.yaml to get your baseline, then wire compare.py into CI so it runs on every PR that touches retrieval, prompts, or the model.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
The problem
Most RAG evaluation is shallow, often relying on a single faithfulness score run once. This fails to catch regressions in retrieval, misses production drift, and ignores the inherent noise in LLM-as-a-judge scoring.
What it does
- Executes a layered evaluation suite covering isolated components (retriever/generator), wired pipelines (RAG Triad), and application-level quality (G-Eval).
- Implements safety gates for toxicity, PII leakage, and system prompt protection to block unsafe deployments.
- Measures operational health including P95 latency, TTFT, cost per query, and reliability rates.
- Provides a noise-aware regression harness that distinguishes real quality drops from standard LLM variance.
- Integrates online production monitoring via LangSmith to detect distribution drift against offline baselines.
Frameworks & tools
Python, DeepEval, LangSmith, and CLI-based automation. Compatible with any RAG stack including LangChain, LlamaIndex, or custom implementations.
Why this beats prompting it yourself
Writing a prompt for 'faithfulness' is easy; building a production-grade test harness that handles claim decomposition, noise-floor calculation, and multi-stage evaluation is complex. This skill provides the architectural spine and statistical rigor needed to trust your RAG metrics.
Use cases
- Setting up a CI/CD gate that blocks RAG deployments if safety or quality metrics drop.
- Benchmarking new embedding models or rerankers against a human-verified golden dataset.
- Monitoring live production traffic for quality drift without requiring golden answers.
- Decomposing RAG failures to identify if the retriever or the generator is the bottleneck.
Known limitations
Requires a clear seam between retriever and generator functions for component-level metrics. LLM-judge scripts require an OpenAI or similar API key for scoring.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 17 days ago
- Passed all security checks, Safe to install
Needs access to