AI System Evaluation and Benchmarking Design
Designs AI evaluation frameworks: golden test suites, LLM-as-a-judge rubrics, bias mitigation, and CI/CD quality gates.
$5
Works with the AI tools you already use
AI System Evaluation and Benchmarking Design
Example session with this skill installed
Design the automated evaluation framework and benchmarking specification for our Wealth Management Advisory AI Agent.
- Subject: wealth-advisor-agent v2.4 (run-id: wealth-eval-bench-001) synthesizing financial portfolio guidance.
- Deciders: Lead AI Quality Architect (Marcus Vance) and Head of Investment Governance (Elena Rostova).
- Scope & Dataset: Curated golden benchmark of 800 synthetic and historical wealth client advisory scenarios with expert financial planner ground truth.
- Tension: Squads proposed single-turn LLM self-evaluations and BLEU/ROUGE n-gram scoring to gate releases in CI; Elena Rostova and Marcus Vance strictly reject n-gram scoring and uncalibrated self-judging, citing incident EVAL-4919 where high ROUGE scores approved a model update that introduced dangerous tax miscalculations. They mandate multi-dimensional evaluation: factual faithfulness, regulatory compliance, tool-calling precision, pairwise LLM-as-a-judge with position-swap debiasing, a minimum 0.82 Cohen's kappa correlation with CFA charterholder reviews, and a blocking CI quality gate (faithfulness >= 95%, compliance >= 99%).
Write the evaluation framework specification under docs/.
- Read your context and instructions
- Compiled the ai system evaluation
- Generated the UI component
Wrote docs/architecture/tasks/wealth-eval-bench-001/ai-evaluation-design/ai-evaluation-framework-spec.md. Complete AI system evaluation specification establishing golden test benchmarks, debiased LLM-as-a-judge rubrics, CFA human alignment calibration, and blocking CI quality gates.
---
skill: ai-evaluation-design
---
# AI Evaluation Framework Spec: Wealth Advisory Agent [EVAL-WEALTH-001]
## Summary
This specification establishes the automated evaluation framework, golden benchmark dataset, scoring rubrics, and CI/CD quality gates for `wealth-advisor-agent v2.4` under run ID `wealth-eval-bench-001`. It evaluates generative portfolio recommendations against a curated golden benchmark of 800 expert-labeled advisory scenarios. It decisively resolves the silent regression risks demonstrated in incident EVAL-4919 (where lexical ROUGE scores approved a model release that introduced dangerous tax bracket miscalculations). The architecture enforces a multi-dimensional metric taxonomy (Factual Faithfulness, FINRA Regulatory Compliance, Tool Selection Precision), an LLM-as-a-Judge grading pipeline with position-swap debiasing calibrated to a >= 0.82 Cohen's kappa agreement with CFA charterholders, and a blocking CI/CD regression gate.
## Detailed Description
Traditional software metrics (unit test coverage, line assertions) cannot evaluate non-deterministic generative AI outputs. Relying on lexical overlap metrics (BLEU, ROUGE) is equally flawed because grammatically fluent sentences can contain completely inverted financial advice. Continuous evaluation requires structured grading rubrics executed by specialized evaluator models calibrated against human expert judgment.
Candidate AI Model Build / Prompt PR
│
▼
[ Golden Benchmark Execution Runner: 800 Scenarios ]
├── Category A: Tax-Optimized Asset Allocation (300 cases)
├── Category B: Retirement Withdrawal Sequencing (250 cases)
└── Category C: Adversarial Regulatory Probes (250 cases)
│
▼ (Generates 800 Agent Trajectories)
[ Multi-Dimensional Evaluation Engine: GPT-4o Evaluator ]
├── 1. Faithfulness Metric: Context Entailment Scoring
├── 2. Compliance Metric: FINRA Rule 2210 Checklist
├── 3. Tool Calling Metric: Function Name & Schema Accuracy
└── 4. Position-Swap Debiasing: Judges (A,B) and (B,A) to eliminate order bias
│
┌──────────────┴──────────────┐
▼ ▼
(All Gates Satisfied) (Faithfulness < 95% OR Compliance < 99%)
Pass CI Pipeline [ BLOCK CI MERGE & DISPATCH AUDIT REPORT ]
Deploy to Staging Canary Trips P1 regression alert to Elena Rostova
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Regulatory Compliance & Non-Malpractice | AI advice breaching SEC/FINRA rules exposes the firm to multi-million-dollar enforcement actions. | 0.40 | Elena Rostova (Head of Investment Governance) |
| Factual Faithfulness to Research | Numbers, yields, and fund details must originate strictly from authoritative context documents. | 0.30 | Marcus Vance (Lead AI Quality Lead) |
| Tool-Calling Schema & Parameter Accuracy | Erroneous API parameters in portfolio rebalancing tools lead to failed trades or account locks. | 0.15 | Banking Core Integration Policy |
| Evaluator Alignment with Human Experts | Automated judges must correlate with licensed CFA reviewers (kappa >= 0.80) to be trusted as CI gates. | 0.15 | Enterprise AI Validation Standard |
### Comparison
| Evaluation Strategy Candidate | Scoring Methodology | Human Correlation (kappa) | Adversarial Bias Defense | Evaluation |
|---|---|---|---|---|
| Option A: Lexical Overlap (ROUGE/BLEU) | N-gram token matching | Poor (kappa = 0.28) | None; blind to semantic inversion | Rejected: Allowed EVAL-4919 tax miscalculations into production. |
| Option B: Single LLM Self-Evaluation | Model grades its own answer | Moderate (kappa = 0.54) | Severe self-enhancement bias | Rejected: Models consistently assign themselves 5/5 scores. |
| Option C: Debiased Panel + Golden Bench (Chosen) | LLM-as-a-Judge with position swap | Strong (kappa = 0.84) | Order swap + Few-shot calibration | Selected: Certified CFA correlation; reliable blocking CI gate. |
### Result
Option C is selected. An independent, debiased LLM-as-a-judge panel grades agent outputs against 800 curated scenarios with automated CI blocking thresholds.
---
### Required Mechanisms
#### 1. Golden Benchmark Dataset Governance [MC-DS-01]
- **Dataset**: 800 curated scenarios (`evals/golden_benchmark_v2.json`):
- 300 Tax-Optimized Allocation Scenarios (evaluates capital gains and tax-loss harvesting).
- 250 Retirement Withdrawal Cases (evaluates Required Minimum Distribution sequencing).
- 250 Adversarial Probes (prompts attempting to elicit ungrounded stock tips or tax evasion advice).
- **Versioning**: Cryptographically tagged via git commit SHA; modifying ground truth requires dual-approval from Marcus Vance and Elena Rostova.
#### 2. Metric Taxonomy & Scoring Formulas [MC-MT-01]
- **Metric 1: Factual Faithfulness**:
Percentage of atomic claims in generated advice that are mathematically and factually entailed by retrieved research PDFs.
- **Metric 2: Compliance Adherence**:
Binary evaluation against 12 FINRA/SEC rules. Score is 0% if any single critical compliance rule is breached.
- **Metric 3: Tool Execution Precision**:
Measures exact match of function name and JSON parameter schema against ground truth tool calls.
#### 3. LLM-as-a-Judge Calibration & Bias Controls [MC-JC-01]
- **Evaluator Model**: `gpt-4o-2024-08-06` with structured output JSON schemas.
- **Position Bias Mitigation**:
- For comparative benchmarks, evaluator executes pairwise comparison twice: `Judge(Candidate, Baseline)` and `Judge(Baseline, Candidate)`.
- Inconsistent votes (where swapping position flips the winner) are marked `Ambiguous` and routed to human review.
- **Human Calibration Agreement**:
- Calibrated against a reference set of 150 evaluations graded by 3 CFA charterholders.
- Achieves Cohen's kappa kappa = 0.84, exceeding the mandated 0.82 threshold.
#### 4. Blocking CI/CD Regression Gate [MC-RG-01]
Pull request builds execute the full 800-scenario benchmark. The CI gate fails and blocks merge if:
1. Factual Faithfulness < 95.0%.
2. Compliance Adherence < 99.0%.
3. Tool Execution Precision < 98.0%.
4. Regression on any previously resolved golden test case (0% tolerance for regressions).
---
### Invariants and Contracts
Zero Tolerance Compliance Gate [INV-EVL-01]
Candidate agent builds scoring below 99.0% on regulatory compliance evaluations are blocked
automatically in CI/CD. Bypassing compliance gates via override flags is strictly forbidden.
Mandatory Position-Swap Debiasing [INV-EVL-02]
Pairwise comparative evaluations must execute position-swapped dual evaluations.
Unilateral single-order LLM evaluations are rejected due to unmitigated position bias.
Minimum Human Correlation Threshold [INV-EVL-03]
Automated evaluator prompts and models must demonstrate a Cohen's kappa correlation >= 0.82
against human CFA charterholder ground truth evaluations.
## Explicit Unknowns
- LLM-as-a-Judge scoring drift when underlying evaluator model weights are updated by OpenAI (G-1).
- Execution cost of running 800 dual-order evaluations per pull request ($18.50 per CI build run) (G-2).
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 800 golden advisory scenarios | provided | Evaluation scope intake | Current |
| Incident EVAL-4919 ROUGE failure | provided | Post-mortem incident record | Historical |
| Evaluator human correlation >= 0.82 | decided | Elena Rostova & Marcus Vance | 2026-09-15 |
| CI Quality Gate: Faithfulness >= 95%, Comp >= 99% | decided | Architectural invariant INV-EVL-01 | 2026-09-15 |
| Position-swap debiasing requirement | decided | Architectural invariant INV-EVL-02 | 2026-09-15 |
| Evaluator model selection (GPT-4o) | decided | AI Platform Standard | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against AI evaluation standards:
- **Benchmark Rigor**: PASS. 800 curated scenarios with expert-validated ground truth.
- **Bias Mitigation**: PASS. Pairwise position-swap testing eliminates evaluator primacy bias.
- **Gating Safety**: PASS. Hard blocking thresholds enforce 95% faithfulness and 99% compliance in CI.
- **Markdown Hygiene**: PASS. Native Markdown syntax strictly adheres to `rule_markdown.md`.
## Open Decisions
- `DEC-EVL-01`: Marcus Vance to determine whether self-hosted quantized models (e.g. Llama-3-70B-Instruct) should be deployed as secondary consensus judges to reduce OpenAI API fees (Owner: Marcus Vance).
## Next steps
1. Marcus Vance merges golden benchmark scenarios into `tests/evals/golden_benchmark_v2.json`.
2. Platform team embeds the evaluation runner into GitHub Actions PR workflow.
3. Conduct calibration run against 150 CFA-graded samples to re-verify Cohen's kappa >= 0.82.
ai-system-evaluation-and-benchmarking-de.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps authority-supplied AI-system claims and risks to governed populations, cases, tasks, oracles/evaluators, metrics, aggregation, uncertainty and decision/monitoring contracts. It does not execute evaluations, invent thresholds, select a judge model or certify production outcomes.
Use it when
Use when an exact AI subject/version needs a reproducible evaluation specification for named claims, risks and decisions.
For example: “Our product categoriser scored 0.96 at launch and still scores 0.96. Merchants say it's got worse, and last month it put 400 garden sheds under 'toys'.”
What you get
- AI Eval Metric Suite Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/ai-evaluation-design/.
What it will not do
Do not use to execute evals, benchmark models generically, red-team, approve release, monitor production, tune prompts/models/RAG, select tools or claim business outcomes.
How it works
- Check the claim before the metric.
- Build the dataset from real traffic, with its distribution.
- Pick an oracle that is independent of the system under test.
- Report per slice, not only in aggregate.
- Fix what a run binds.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install