Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    incident derived agent evaluation dataset

    1

    Construct provenance-linked, deduplicated evaluation datasets from real-world agent incident logs.

    $15

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    incident derived agent evaluation dataset

    incident derived agent evaluation dataset

    Example session with this skill installed

    Build an evaluation dataset from the three redacted API timeout incidents in incident-logs/q3-failures.json. Deduplicate and split by incident family.

    • Read your context and instructions
    • Compiled the incident derived agent

    I have processed the incident logs, deduplicated the failure traces, and generated a provenance-linked dataset manifest. The incidents were grouped into the 'Connection Timeout' family and split to preserve a held-out set.

    Incident-Derived Dataset Manifest

    MetricValue
    Total Incidents Processed3
    Unique Prompt Families1
    Deduplicated Records2
    Source Provenanceincident-logs/q3-failures.json
    Split Strategy50% Production / 50% Held-out
    Artifact Path./outputs/dataset_2026_09_18.json

    Dataset Content Summary

    [
      {
        "id": "CASE-001",
        "incident_family": "api_timeout_retry_exhaustion",
        "minimal_context": "User requests high-latency reporting endpoint; auth token expires during 30s wait.",
        "expected_behavior": "Agent should refresh token or return 408 status; current failure shows 500 loop.",
        "provenance_ref": "q3-failures.json:line_45",
        "tags": ["network", "auth-lifecycle"]
      },
      {
        "id": "CASE-002",
        "incident_family": "api_timeout_retry_exhaustion",
        "minimal_context": "Parallel execution of 5 data fetch calls exceeds concurrent connection limit.",
        "expected_behavior": "Queue management or exponential backoff implementation.",
        "provenance_ref": "q3-failures.json:line_112",
        "tags": ["concurrency", "rate-limiting"]
      }
    ]
    

    Next steps

    1. Run scripts/build_dataset.py to export the deterministic JSONL for your evaluation pipeline.
    2. Review the semantic duplicates in ./outputs/review_queue.md to confirm the deduplication logic.
    3. Apply the held-out set to your CI/CD gate for the next agent release.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Extract reproducible failure prompts from production logs.Deduplicate incident reports to prevent test set contamination.Link evaluation cases to specific source provenance and versions.Isolate fact-based ground truth from speculative incident explanations.

    About this skill

    The problem

    General LLM benchmarks fail to capture how your specific agent breaks in production. Generic evaluation datasets lack the provenance needed to trace failures back to real-world incident root causes.

    What it does

    • Extracts minimal reproducible prompts from redacted incident logs and failure traces.
    • Groups paraphrases and near-duplicate incidents to prevent data leakage across splits.
    • Separates verified facts from proposed explanations to ensure objective ground truth.
    • Generates a deterministic dataset manifest with linked source provenance and versioning.
    • Identifies label conflicts and semantic duplicates for manual review.

    Frameworks & tools

    Python scripts (build_dataset.py), JSON schema for incident fixtures, and provenance-linked reporting.

    Why this beats prompting it yourself

    Manual dataset creation often leads to data contamination where similar prompts appear in both train and test sets. This skill enforces strict evidence discipline, ensuring every evaluation case is linked to a real incident without leaking sensitive credentials or model-generated hallucinations.

    Use cases

    • Building a gold-standard evaluation set for agentic workflows.
    • Deduplicating incident reports to create clean fine-tuning data.
    • Auditing agent performance against historical production failures.
    • Preparing held-out test sets for rigorous model deployment gates.

    Known limitations

    Requires pre-redacted incident cases as input. Does not perform model evaluation itself or invent ground truth where source data is missing.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 13 days ago

    • Passed all security checks, Safe to install

    Needs access to

    Json Schema

    Listed13 days ago

    What's inside

    Frequently Asked Questions