More screenshots
Works with the AI tools you already use
Research Evaluation Designer
Design an evidence-backed evaluation rubric for LLM, RAG, research, or agent outputs.
$19.99
Research Evaluation Designer
Example session with this skill installed
Design a HYBRID evaluation rubric for a clinical-QA RAG system in the medical domain. Include the RAGAS core, domain-specific clinical safety dimensions, anchored human review, a HIPAA-related evaluation overlay, and a 20–30 example pilot. Do not present any threshold as production-validated without pilot evidence.
- Read your context and instructions
- Compiled the research evaluation designer
Clinical-QA RAG Evaluation Blueprint
System: clinical-QA RAG pipeline | Mode: HYBRID | Domain: medical
| Dimension | Method | Initial threshold | Evidence status |
|---|---|---|---|
| Faithfulness | Claim extraction plus NLI against retrieved context | 0.80 | Pilot target, not production validated |
| Answer relevance | Reverse-question generation plus cosine similarity | 0.80 | Pilot target |
| Context precision | Useful chunks divided by total retrieved chunks | 0.70 | Pilot target |
| Context recall | Ground-truth statements covered divided by total statements | 0.80 | Pilot target |
| Clinical accuracy | Claims checked against a curated medical-guidelines set | 0.85 | Domain-specific proposed metric |
| Contraindication detection | Dangerous recommendations flagged across seeded cases | 0.90 | Requires explicit pilot evidence before adoption |
Human anchors
Coherence: 1 = contradictory or illogical; 3 = mostly coherent with minor gaps; 5 = fully coherent clinical reasoning.
Completeness: 1 = misses most sub-questions; 3 = addresses the main question only; 5 = addresses every stated sub-question.
Compliance overlay
For the supplied HIPAA requirement, verify that the evaluation set contains no patient identifiers and retain provenance for clinical claims. This mapping supports evaluation design; it does not establish legal compliance.
Pilot
Use 20–30 stratified examples including retrieval misses, irrelevant context, contraindication traps, and complete answers. Double-rate a subset, examine disagreement, revise anchors, and recalibrate thresholds before any production gate.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Design a domain-specific evaluation rubric for LLM, RAG, research, or multi-agent outputs. Research Evaluation Designer selects evidence-backed dimensions, combines automated metrics with anchored human scales, adds compliance mappings only when supplied, rejects unsupported thresholds, and includes a calibration plan without pretending the rubric is already validated.
You receive a dimension table, measurement methods, source register, human-rating anchors, compliance overlay, threshold status, and a pilot plan. The package includes the full SKILL.md workflow, RAGAS references, calibration guidance, compliance matrix, JSON rubric schema, deterministic checker, worked sample, tests, and a six-page PDF buyer guide.
It designs rubrics; it does not run evaluations, generate a dataset, establish regulatory compliance, or claim production validity without a pilot.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 1 month ago
- Passed all security checks, Safe to install