AI Evaluation and Benchmark Platform Architect

    1

    Architects enterprise AI evaluation systems: offline golden benchmarks, online judge fleets, red teaming, and CI gates.

    $12

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    AI Evaluation and Benchmark Platform Architect

    Example session with this skill installed

    Architect the enterprise AI evaluation platform and continuous benchmark governance for our FinTech AI Suite.

    • Subject: FinTech AI Evaluation Platform (run-id: fintech-ai-eval-001) evaluating 25 generative AI applications.
    • Deciders: Principal AI Quality Architect (Marcus Vance) and Chief AI Officer (Elena Rostova).
    • Scale: Supporting 180 AI engineers; evaluating 120,000 daily production transactions; 4,500 curated golden test cases.
    • Tension: Squads proposed decentralized ad-hoc Python scripts with unversioned prompt rubrics and zero online monitoring; Elena Rostova and Marcus Vance strictly reject un-governed scripts, citing incident EVAL-5102 where unmonitored model drift caused a 14% drop in loan underwriting accuracy over 6 weeks. They mandate an enterprise platform: centralized golden dataset lifecycle management, an automated multi-judge consensus fleet (cross-evaluating with GPT-4o and Claude 3.5 Sonnet), continuous online shadow evaluation (sampling 5% of production inference), adversarial red-teaming harnesses, and blocking CI/CD release gates.

    Write the evaluation platform architecture overview under docs/.

    • Read your context and instructions
    • Compiled the ai evaluation

    Wrote docs/architecture/evaluation-architect/00-overview/evaluation-architect-overview.md. Complete AI evaluation platform architecture specifying centralized benchmark datasets, multi-judge consensus fleets, online shadow evaluation, and automated CI/CD quality gates.

    ---
    skill: evaluation-architect
    ---
    
    # AI Evaluation Platform Architecture: FinTech AI Suite [EVAL-PLAT-001]
    
    ## Summary
    
    This specification establishes the enterprise AI evaluation platform, benchmark governance, and continuous model validation architecture for the FinTech AI Suite under run ID `fintech-ai-eval-001`. It governs 25 generative AI applications across wealth advisory, fraud investigation, and credit underwriting, supporting 180 AI engineers and evaluating 120,000 daily production interactions. It decisively resolves the silent performance degradation demonstrated in incident EVAL-5102 (where unmonitored model weight drift eroded credit underwriting accuracy by 14% over 6 weeks without detection). The architecture enforces a unified evaluation control plane: centralized golden benchmark dataset versioning, an automated multi-judge consensus fleet (GPT-4o and Claude 3.5 Sonnet with position-swap debiasing), continuous 5% online production shadow evaluation, automated adversarial red-teaming, and blocking CI/CD deployment gates.
    
    ## Detailed Description
    
    Generative AI applications cannot be reliably governed through ad-hoc developer scripts or isolated test runs. Without a centralized evaluation platform, squads use inconsistent grading rubrics, ground-truth datasets become stale, and silent production drift goes undetected until customers report business errors.
    
    

    Model Development & Production Inference Streams
    │
    ┌────────────────┴────────────────┐
    ▼ (Pre-Deployment Phase) ▼ (Production Phase)
    [ Offline Golden Benchmark Runner ] [ Online Production Shadow Stream (5%) ]
    ├── 4,500 Versioned Test Scenarios ├── Async Kafka Mirror: 6,000 calls/day
    └── Automated Adversarial Red Teaming └── Anonymized PII Sanitization Pipe
    │ │
    └──────────────────┬────────────────────┘
    ▼
    [ Multi-Judge Consensus Evaluation Engine ]
    ├── Judge A: GPT-4o (Structured Schema JSON Rubric)
    ├── Judge B: Claude 3.5 Sonnet (Position-Swapped Evaluation)
    └── Disagreement Resolution: Kappa >= 0.85; routes delta to Human SME Queue
    │
    ┌───────────────────────────┴───────────────────────────┐
    ▼ ▼
    (Pre-Merge Quality Gate: Passed) (Drift Detected: Score Drops > 3%)
    Authorize Staging / Canary Deploy [ Automated Mitigation Tripwire ]
    ├── Halt release pipeline
    └── Alert Marcus Vance & Elena Rostova

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Continuous Drift Detection & Early Warning | Silent model degradation directly threatens credit risk models and financial underwriting (EVAL-5102). | 0.35 | Elena Rostova (Chief AI Officer) |
    | Evaluator Objectivity & Bias Elimination | Single-model judges exhibit severe self-enhancement and position bias, invalidating benchmark results. | 0.30 | Marcus Vance (Lead Quality Architect) |
    | Standardized Golden Dataset Governance | Centralized, cryptographically versioned test sets prevent squads from gaming evaluation metrics. | 0.20 | Enterprise AI Governance Board |
    | Developer Velocity & CI Feedback (< 25 min) | Running 4,500 benchmark evaluations must execute rapidly in CI to prevent PR merge bottlenecks. | 0.15 | AI Engineering Delivery SLA |
    
    
    ### Comparison
    
    | Evaluation Architecture Candidate | Evaluation Topology | Judge Consensus Model | Online Drift Tracking | Evaluation |
    |---|---|---|---|---|
    | Option A: Decentralized Squad Scripts | Ad-hoc local Python scripts | Single self-selected LLM | None (Zero production monitoring) | Rejected: Caused EVAL-5102 underwriting drift disaster. |
    | Option B: Commercial Evaluation SaaS | Hosted external evaluation cloud | Proprietary black-box judge | Batch log upload | Rejected: Exposes confidential customer banking data to third-party SaaS. |
    | Option C: Centralized Platform Engine (Chosen) | Internal K8s evaluation plane | Dual-Judge Consensus (GPT-4o + Sonnet) | Continuous 5% online shadow stream | Selected: Complete data sovereignty, debiased consensus, zero drift. |
    
    
    ### Result
    
    Option C is selected. An enterprise evaluation control plane coordinates golden datasets, dual-judge consensus grading, and online production drift monitoring.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Centralized Benchmark Dataset Lifecycle [MC-DS-01]
    - **Storage & Registry**: MLflow / DVC dataset registry backed by immutable S3 storage (`s3://bank-ai-benchmarks/`).
    - **Benchmark Composition**: 4,500 golden scenarios partitioned into 3 tiers:
      - Tier 1: Core Domain Capabilities (2,500 scenarios across underwriting, wealth, support).
      - Tier 2: Edge Cases & Multi-Step Reasoning (1,000 complex dispute scenarios).
      - Tier 3: Adversarial Red-Teaming (1,000 jailbreak and prompt-injection probes).
    - **Versioning**: Datasets are immutable and tagged with cryptographic SHA-256 commit hashes. Updating ground truth requires dual-approval.
    
    #### 2. Multi-Judge Consensus Evaluation Engine [MC-JE-01]
    - **Dual-Model Panel**: Evaluates outputs using two independent model families:
      - Judge 1: `gpt-4o-2024-08-06`
      - Judge 2: `claude-3-5-sonnet-20241022`
    - **Consensus & Debiasing Protocol**:
      - Pairwise evaluations execute with position swapping `(A, B)` and `(B, A)` to eliminate order bias.
      - If Judge 1 and Judge 2 agree, score commits to telemetry.
      - If judges disagree, the sample is automatically escalated to the **Human Expert Review Queue** (CFA / underwriter review).
    
    #### 3. Continuous Online Shadow Evaluation [MC-SE-01]
    - 5% of all live production inference requests (6,000 calls/day) are sampled out-of-band via Kafka mirror.
    - PII sanitization filters redact account numbers and customer names.
    - Evaluator fleet grades production outputs continuously across 24-hour sliding windows.
    - **Drift Tripwire**: If 24-hour moving average quality score drops by >= 3.0% compared to baseline, an automated P1 incident alerts Elena Rostova.
    
    #### 4. Automated CI/CD Deployment Gates [MC-DG-01]
    - Every candidate model commit or prompt update must execute the full 4,500 golden benchmark suite in GitHub Actions (concurrency: 50 parallel workers; execution time: 18 minutes).
    - **Passing Thresholds**:
      - Domain Accuracy >= 94.0%.
      - Regulatory Compliance Score >= 99.5%.
      - Adversarial Jailbreak Defense >= 99.8%.
      - Regression on previously passing Tier 1 test cases: **0 allowed**.
    
    ---
    
    ### Invariants and Contracts
    
        Mandatory Dual-Judge Consensus Invariant [INV-EVP-01]
          Evaluations determining deployment gating must use at least two distinct model families.
          Single-model evaluations or self-evaluations (model grading itself) fail platform gate checks.
    
        Zero Regression on Tier-1 Benchmarks [INV-EVP-02]
          Candidate AI models must achieve zero regressions on previously certified Tier 1 golden scenarios.
          Any regression on a core capability benchmark aborts automated deployment pipelines.
    
        Continuous Production Drift Monitoring Mandate [INV-EVP-03]
          Production generative AI applications must sample at least 5% of live transactions for continuous
          evaluation. Running unmonitored production generative workloads is blocked by platform policy.
    
    ## Explicit Unknowns
    
    - Token API evaluation billing costs when scaling benchmark suite from 4,500 to 20,000 scenarios (G-1).
    - Human SME queue turnaround times when complex financial dispute cases require manual review (G-2).
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | 25 generative AI applications | provided | Scope intake | Current |
    | 120,000 daily production transactions | provided | Operational volume intake | Current |
    | 4,500 curated golden benchmark scenarios | provided | Dataset intake | Current |
    | Incident EVAL-5102 14% underwriting drift | provided | Post-mortem evidence | Historical |
    | Dual-judge consensus (GPT-4o + Sonnet) | decided | Marcus Vance (Lead AI Architect) | 2026-09-15 |
    | 5% online production shadow sampling | decided | Elena Rostova (Chief AI Officer) | 2026-09-15 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against AI evaluation platform standards:
    - **Platform Scope**: PASS. Covers offline golden benchmarks, online shadow evaluation, and red-teaming.
    - **Judge Governance**: PASS. Dual-model consensus with position-swap debiasing eliminates single-judge bias.
    - **Drift Safety**: PASS. Continuous 5% sampling detects production degradation within 24 hours.
    - **Markdown Hygiene**: PASS. Native Markdown syntax strictly adheres to `rule_markdown.md`.
    
    ## Open Decisions
    
    - `DEC-EVP-01`: Elena Rostova to determine whether human SME reviewers should be incentivized with platform bounty credits for resolving ambiguous evaluation disputes (Owner: Elena Rostova).
    
    ## Next steps
    
    1. Marcus Vance provisions central evaluation worker cluster on AWS EKS with Ray distributed computing.
    2. Platform team establishes DVC repository for versioned golden benchmark datasets.
    3. Deploy continuous 5% shadow sampling pipeline on the underwriting service to establish live baseline telemetry.
    

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define golden datasets and oracle judges for RAG system validationMap product claims to specific metrics and failure mode slicesEstablish CI gates and offline-to-online evaluation lifecyclesDesign red teaming protocols and human reviewer calibration workflows

    About this skill

    What it does

    This skill owns the evaluation decision system that connects what an AI system is allowed to claim and what can go wrong to representative tasks, datasets, evaluators, baselines, slices, uncertainty, lifecycle evidence and policy-owned decisions. It designs how evidence is produced, interpreted, invalidated and handed to release or risk owners; it does not manufacture favorable scores or decide policy by itself.

    Use it when

    • Product/model/RAG/agent claims and failure risks need explicit evaluation coverage
    • Tasks, trajectory levels, environment state and observable outcomes must be modeled
    • Dataset population, lineage, splits, consent/license, freshness and access require governance
    • Deterministic checks, statistical metrics, model judges and human reviewers must be combined
    • Baselines/comparators and equivalent run conditions need definition
    • Critical populations, risks or failure modes require named slices rather than aggregate-only results

    For example: “We want to know if the new retrieval config is better before we ship it. The team has been eyeballing twenty examples.”

    What you get

    • architecture/evaluation-architect/README.md
    • architecture/evaluation-architect/00-overview/evaluation-architect-overview.md
    • architecture/evaluation-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to run an existing benchmark, score one output, write one judge prompt or test, clean or label a dataset, implement a harness/dashboard/tracing pipeline, tune a model/prompt/RAG system, compare vendor leaderboards, audit only failure traces, or issue a release verdict without decision authority because words such as eval, quality, hallucination, benchmark, metric, score, judge, regression, test, or leaderboard appear.

    How it works

    1. Write down the claim.
    2. Name what going wrong looks like.
    3. Pick the oracle before the dataset.
    4. Bound the system under test.
    5. Separate offline from online.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions