Works with the AI tools you already use
incident derived agent evaluation dataset
Construct provenance-linked, deduplicated evaluation datasets from real-world agent incident logs.
$15
incident derived agent evaluation dataset
Example session with this skill installed
Build an evaluation dataset from the three redacted API timeout incidents in incident-logs/q3-failures.json. Deduplicate and split by incident family.
- Read your context and instructions
- Compiled the incident derived agent
I have processed the incident logs, deduplicated the failure traces, and generated a provenance-linked dataset manifest. The incidents were grouped into the 'Connection Timeout' family and split to preserve a held-out set.
Incident-Derived Dataset Manifest
| Metric | Value |
|---|---|
| Total Incidents Processed | 3 |
| Unique Prompt Families | 1 |
| Deduplicated Records | 2 |
| Source Provenance | incident-logs/q3-failures.json |
| Split Strategy | 50% Production / 50% Held-out |
| Artifact Path | ./outputs/dataset_2026_09_18.json |
Dataset Content Summary
[
{
"id": "CASE-001",
"incident_family": "api_timeout_retry_exhaustion",
"minimal_context": "User requests high-latency reporting endpoint; auth token expires during 30s wait.",
"expected_behavior": "Agent should refresh token or return 408 status; current failure shows 500 loop.",
"provenance_ref": "q3-failures.json:line_45",
"tags": ["network", "auth-lifecycle"]
},
{
"id": "CASE-002",
"incident_family": "api_timeout_retry_exhaustion",
"minimal_context": "Parallel execution of 5 data fetch calls exceeds concurrent connection limit.",
"expected_behavior": "Queue management or exponential backoff implementation.",
"provenance_ref": "q3-failures.json:line_112",
"tags": ["concurrency", "rate-limiting"]
}
]
Next steps
- Run
scripts/build_dataset.pyto export the deterministic JSONL for your evaluation pipeline. - Review the semantic duplicates in
./outputs/review_queue.mdto confirm the deduplication logic. - Apply the held-out set to your CI/CD gate for the next agent release.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
The problem
General LLM benchmarks fail to capture how your specific agent breaks in production. Generic evaluation datasets lack the provenance needed to trace failures back to real-world incident root causes.
What it does
- Extracts minimal reproducible prompts from redacted incident logs and failure traces.
- Groups paraphrases and near-duplicate incidents to prevent data leakage across splits.
- Separates verified facts from proposed explanations to ensure objective ground truth.
- Generates a deterministic dataset manifest with linked source provenance and versioning.
- Identifies label conflicts and semantic duplicates for manual review.
Frameworks & tools
Python scripts (build_dataset.py), JSON schema for incident fixtures, and provenance-linked reporting.
Why this beats prompting it yourself
Manual dataset creation often leads to data contamination where similar prompts appear in both train and test sets. This skill enforces strict evidence discipline, ensuring every evaluation case is linked to a real incident without leaking sensitive credentials or model-generated hallucinations.
Use cases
- Building a gold-standard evaluation set for agentic workflows.
- Deduplicating incident reports to create clean fine-tuning data.
- Auditing agent performance against historical production failures.
- Preparing held-out test sets for rigorous model deployment gates.
Known limitations
Requires pre-redacted incident cases as input. Does not perform model evaluation itself or invent ground truth where source data is missing.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 13 days ago
- Passed all security checks, Safe to install
Needs access to