- Home
- Skills
- Prompt & Skill Engineering
- Benchmark Forge
Works with the AI tools you already use
Benchmark Forge
Turns AI prompts and agent behaviors into rigorous, diagnostic evaluation datasets that expose failure modes, edge cases, inconsistencies, and behavioral weaknesses.
$12
Benchmark Forge
Example session with this skill installed
I’m building a customer-support ticket classifier.
SYSTEM PROMPT
You classify incoming customer-support messages into exactly one of these categories:
- BILLING — charges, invoices, refunds, payments, subscriptions, or pricing
- TECHNICAL — bugs, errors, crashes, login problems, or product malfunctions
- ACCOUNT — account settings, profile changes, account access, or account management
- SHIPPING — delivery status, shipping delays, tracking, or damaged/missing packages
- CANCELLATION — requests to cancel a subscription or service
Return ONLY the category name. Do not explain your answer. If a message seems to belong to multiple categories, choose the category representing the customer's primary intent.
Generate a diagnostic evaluation dataset of 25 cases for this classifier.
I especially want to know whether the prompt handles ambiguous tickets, overlapping categories, irrelevant details, and small wording changes consistently.
- Read your context and instructions
- Compiled the benchmark forge
- Generated the document
This classifier is small (5 fixed labels) but its own definitions already create fault lines worth probing — most notably that "login problems" is explicitly TECHNICAL while "account access" is explicitly ACCOUNT, and there's no catch-all for out-of-taxonomy tickets. I designed the set around those structural weak points plus the ambiguity/overlap/wording-sensitivity axes you asked about.
Coverage: 5 canonical baseline · 3 contradiction (TECHNICAL vs ACCOUNT boundary) · 4 overlapping-intent · 2 genuinely ambiguous · 2 distractor · 4 perturbation (2 pairs) · 2 contrastive (1 pair) · 2 out-of-taxonomy · 1 stress/compositional · 1 instruction-adherence-under-pressure.The three most diagnostic cases are B1–B3 — they probe a contradiction baked into your own prompt (login failures are TECHNICAL, account access is ACCOUNT, but real tickets blur the two constantly). How the model handles those, plus O1/O4/S1 (explicit action vs. underlying grievance), will tell you more than raw accuracy will.
Also worth noting: N1/N2 aren't really testable in the normal sense — there's no correct label for them under the current 5-category taxonomy, so they're there to quantify how often you'd be force-fitting real traffic rather than to grade the model.
Support Ticket Classifier — Diagnostic Eval Set v1
25 cases targeting: ambiguity handling, category overlap/primary-intent resolution, distractor robustness, phrasing sensitivity, contrastive sensitivity, out-of-taxonomy behavior, and instruction adherence (label-only output) under pressure.
Categories: BILLING · TECHNICAL · ACCOUNT · SHIPPING · CANCELLATION
1. Canonical baseline
| ID | Input | Expected | Notes |
|---|---|---|---|
| C1 | "I was charged twice for my subscription this month, can I get a refund for the duplicate charge?" | BILLING | Clean sanity check |
| C2 | "The app crashes every time I try to upload a photo." | TECHNICAL | Clean sanity check |
| C3 | "I want to update my email address on file and change my display name." | ACCOUNT | Clean sanity check |
| C4 | "My package says delivered but I never received it." | SHIPPING | Clean sanity check |
| C5 | "Please cancel my subscription effective immediately." | CANCELLATION | Clean sanity check |
2. Contradiction: TECHNICAL ("login problems") vs ACCOUNT ("account access")
The prompt itself puts these two phrases in different buckets even though they describe the same real-world event. This is the single highest-value place to test.
| ID | Input | Expected | Rationale |
|---|---|---|---|
| B1 | "I can't log into my account, it says wrong password every time even after reset." | TECHNICAL | Framed as a malfunction (reset not working) — matches "login problems" literally |
| B2 | "I'm locked out of my account and need help regaining access." | ACCOUNT | Framed as an access/administrative request, no malfunction implied — matches "account access" literally |
| B3 | "I forgot my password and the reset link isn't working, so I can't get into my account settings to update my billing info." | TECHNICAL | Compositional near-miss of B1/B2 plus a BILLING lure — the blocking failure (broken reset link) is the actual problem; billing/account are downstream goals the customer can't reach yet |
What this pair tests: whether the model resolves the spec's internal ambiguity consistently (B1 and B2 should not both get the same label just because both mention "account") or whether it pattern-matches on the word "account" regardless of framing.
3. Overlapping categories — primary intent resolution
| ID | Input | Expected | Rationale |
|---|---|---|---|
| O1 | "I've been overcharged three times now, I'm done — cancel my account and refund the money." | CANCELLATION | Explicit cancellation action requested, even though grievance is billing |
| O2 | "My order never arrived and it's been 3 weeks, I want my money back." | SHIPPING (defensible: BILLING) | Genuinely contestable — root cause is non-delivery, but the stated ask is a refund. Flag for human calibration rather than treating as a hard pass/fail |
| O3 | "Every time I try to update my payment method the page just shows a 500 error." | TECHNICAL | Symptom is a malfunction; the billing goal shouldn't override an explicit bug report |
| O4 | "The app has been broken for two weeks and nobody's fixed it, I want to cancel." | CANCELLATION | Explicit action requested (cancel) outweighs the technical grievance behind it |
What this tests: consistency of the "primary intent" rule — O1/O4 share a pattern (explicit action + unrelated grievance → action wins) that a good classifier should apply uniformly; O3 tests whether a bug report gets contaminated by its billing context; O2 has no clean answer and should be used to check evaluator/rubric agreement, not scored as strict pass/fail.
4. Genuinely ambiguous
| ID | Input | Expected | Rationale |
|---|---|---|---|
| A1 | "My subscription isn't working right." | No single defensible label — could be BILLING (payment failed), TECHNICAL (feature broken), or ACCOUNT (wrong plan) | Tests whether the model invents unstated specificity to force a confident answer, or defaults to a majority-class guess |
| A2 | "Can you help me with my account?" | No single defensible label | Contains the keyword "account" but no actionable request — tests keyword-matching vs. genuine intent detection |
What this tests: since the system prompt forces exactly one label with no "insufficient information" option, these cases reveal whether the model's confidence is calibrated (same vague input classified differently across runs would indicate guessing) rather than whether it picks the "correct" label.
5. Distractors (irrelevant but realistic detail)
| ID | Input | Expected | Rationale |
|---|---|---|---|
| D1 | "Hi! Hope you're having a great day. I've been a customer for 5 years and love your product, my dog also loves chewing on the packaging haha. Anyway, I noticed my card was charged $49.99 but I'm on the $29.99 plan, can you fix that?" | BILLING | Chit-chat and small talk should not dilute the clear billing ask |
| D2 | "I ordered a new charger cable to replace the one that's not working, but it's now 2 weeks late and I still don't have it." | SHIPPING | "not working" describes the old cable (background), not a request — the actual ask is about the late order |
6. Perturbation pairs (should NOT flip label — semantic invariance)
| ID | Input | Expected | Pair |
|---|---|---|---|
| P1 | "I want to cancel my subscription." | CANCELLATION | — |
| P1' | "Please terminate my plan, I no longer wish to be billed." | CANCELLATION | Same intent as P1, different wording/length/formality |
| P2 | "Where is my order?" | SHIPPING | — |
| P2' | "I placed an order on the 3rd and it still hasn't shipped, what's going on?" | SHIPPING | Same intent as P2, more specific and less terse |
What this tests: if P1/P1' or P2/P2' diverge, the model is phrasing-sensitive rather than intent-sensitive.
7. Contrastive pair (SHOULD flip label — semantic sensitivity)
| ID | Input | Expected | Pair |
|---|---|---|---|
| X1 | "I want to cancel my subscription." | CANCELLATION | — |
| X1' | "I want to downgrade my subscription." | BILLING (plan/pricing change, not termination) | One-word substitution that should flip the result — tests whether "cancel"-adjacent language over-triggers CANCELLATION |
8. Out-of-taxonomy (no category actually fits)
| ID | Input | Expected | Rationale |
|---|---|---|---|
| N1 | "Do you guys have a referral program? I'd love to invite my friends." | No correct label exists | Sales/marketing inquiry — none of the 5 categories apply. Since the prompt has no fallback/"other" option, whatever label is returned is a forced misclassification by construction. The finding here isn't "which label" but "how often real traffic will contain input the taxonomy can't represent" |
| N2 | "Just wanted to say thank you, your support team fixed my issue so fast!" | No correct label exists | Pure compliment, no actionable request in any category |
What this tests: this is a taxonomy-completeness gap, not a model failure — worth flagging to you directly since no amount of prompt-tuning fixes it without adding a catch-all category (e.g. OTHER/GENERAL).
9. Stress / compositional
| ID | Input | Expected | Rationale |
|---|---|---|---|
| S1 | "So this has been a nightmare. First my package showed up damaged, then when I tried to report it the app kept crashing, then I got charged again for a 'replacement' I never approved, and now I just want to cancel everything and never deal with this again." | CANCELLATION | Four category-relevant threads (SHIPPING, TECHNICAL, BILLING, and an explicit cancel request) in one message — the explicit terminal action ("cancel everything") should still win under the primary-intent rule, same pattern as O1/O4 |
10. Instruction adherence under pressure
| ID | Input | Expected | Rationale |
|---|---|---|---|
| F1 | "Can you tell me which category this is and why: 'My package arrived damaged and I want a replacement.'" | SHIPPING (label only, no explanation) | The embedded ticket text explicitly asks for an explanation — tests whether the "return ONLY the category name, do not explain" instruction holds even when the input tries to elicit one |
Structural findings independent of any single test result
Built-in category collision: "login problems" (TECHNICAL) and "account access" (ACCOUNT) describe overlapping real-world events. Section 2 cases (B1–B3) will reveal whether the model has a consistent implicit rule for this, but the fix is really at the prompt level — consider adding a tiebreaker clause (e.g. "if login/access issues stem from a technical error, use TECHNICAL; if the account itself is locked/restricted with no error, use ACCOUNT").
2.
No fallback category: N1/N2 will always be forced into a wrong label. If off-taxonomy tickets are common in production, add an OTHER/GENERAL bucket or a confidence/abstention mechanism rather than relying on eval results to catch this.
3.
"Primary intent" is undefined: the prompt says to choose the category representing "primary intent" but gives no rule for when an explicit action (cancel/refund) competes with an underlying grievance (bug/shipping/billing). Cases O1–O4 and S1 test whether the model has converged on a consistent implicit rule ("explicit requested action beats background grievance") — if it has, consider codifying that rule explicitly in the prompt so it's not left to inference.
benchmark-forge.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Why buy it?
Because manually designing good AI evaluations is surprisingly difficult. Generating examples is easy; knowing which examples are capable of revealing meaningful failures is the hard part. A weak test suite can give you false confidence by containing mostly obvious cases. This skill is designed to uncover the cases that ordinary prompting tends to miss — the boundary conditions, conflicting requirements, subtle semantic changes, misleading context, incomplete information, and realistic failure scenarios where seemingly capable models break down. Instead of asking an AI system, “Can you handle this example?”, you can use this skill to systematically ask: Where does the capability stop working? Why does it stop working? What evidence would reveal that failure? And can we reliably test it again after the system changes? That is the difference between having a collection of examples and having an actual AI evaluation system.How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 20 days ago
- Passed all security checks, Safe to install