Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    Benchmark Forge

    1

    Turns AI prompts and agent behaviors into rigorous, diagnostic evaluation datasets that expose failure modes, edge cases, inconsistencies, and behavioral weaknesses.

    $12

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    Benchmark Forge

    Benchmark Forge

    Example session with this skill installed

    I’m building a customer-support ticket classifier.

    SYSTEM PROMPT

    You classify incoming customer-support messages into exactly one of these categories:

    • BILLING — charges, invoices, refunds, payments, subscriptions, or pricing
    • TECHNICAL — bugs, errors, crashes, login problems, or product malfunctions
    • ACCOUNT — account settings, profile changes, account access, or account management
    • SHIPPING — delivery status, shipping delays, tracking, or damaged/missing packages
    • CANCELLATION — requests to cancel a subscription or service

    Return ONLY the category name. Do not explain your answer. If a message seems to belong to multiple categories, choose the category representing the customer's primary intent.

    Generate a diagnostic evaluation dataset of 25 cases for this classifier.

    I especially want to know whether the prompt handles ambiguous tickets, overlapping categories, irrelevant details, and small wording changes consistently.

    • Read your context and instructions
    • Compiled the benchmark forge
    • Generated the document

    This classifier is small (5 fixed labels) but its own definitions already create fault lines worth probing — most notably that "login problems" is explicitly TECHNICAL while "account access" is explicitly ACCOUNT, and there's no catch-all for out-of-taxonomy tickets. I designed the set around those structural weak points plus the ambiguity/overlap/wording-sensitivity axes you asked about.

    Coverage: 5 canonical baseline · 3 contradiction (TECHNICAL vs ACCOUNT boundary) · 4 overlapping-intent · 2 genuinely ambiguous · 2 distractor · 4 perturbation (2 pairs) · 2 contrastive (1 pair) · 2 out-of-taxonomy · 1 stress/compositional · 1 instruction-adherence-under-pressure.The three most diagnostic cases are B1–B3 — they probe a contradiction baked into your own prompt (login failures are TECHNICAL, account access is ACCOUNT, but real tickets blur the two constantly). How the model handles those, plus O1/O4/S1 (explicit action vs. underlying grievance), will tell you more than raw accuracy will.

    Also worth noting: N1/N2 aren't really testable in the normal sense — there's no correct label for them under the current 5-category taxonomy, so they're there to quantify how often you'd be force-fitting real traffic rather than to grade the model.

    Support Ticket Classifier — Diagnostic Eval Set v1

    25 cases targeting: ambiguity handling, category overlap/primary-intent resolution, distractor robustness, phrasing sensitivity, contrastive sensitivity, out-of-taxonomy behavior, and instruction adherence (label-only output) under pressure.

    Categories: BILLING · TECHNICAL · ACCOUNT · SHIPPING · CANCELLATION


    1. Canonical baseline

    IDInputExpectedNotes
    C1"I was charged twice for my subscription this month, can I get a refund for the duplicate charge?"BILLINGClean sanity check
    C2"The app crashes every time I try to upload a photo."TECHNICALClean sanity check
    C3"I want to update my email address on file and change my display name."ACCOUNTClean sanity check
    C4"My package says delivered but I never received it."SHIPPINGClean sanity check
    C5"Please cancel my subscription effective immediately."CANCELLATIONClean sanity check

    2. Contradiction: TECHNICAL ("login problems") vs ACCOUNT ("account access")

    The prompt itself puts these two phrases in different buckets even though they describe the same real-world event. This is the single highest-value place to test.

    IDInputExpectedRationale
    B1"I can't log into my account, it says wrong password every time even after reset."TECHNICALFramed as a malfunction (reset not working) — matches "login problems" literally
    B2"I'm locked out of my account and need help regaining access."ACCOUNTFramed as an access/administrative request, no malfunction implied — matches "account access" literally
    B3"I forgot my password and the reset link isn't working, so I can't get into my account settings to update my billing info."TECHNICALCompositional near-miss of B1/B2 plus a BILLING lure — the blocking failure (broken reset link) is the actual problem; billing/account are downstream goals the customer can't reach yet

    What this pair tests: whether the model resolves the spec's internal ambiguity consistently (B1 and B2 should not both get the same label just because both mention "account") or whether it pattern-matches on the word "account" regardless of framing.

    3. Overlapping categories — primary intent resolution

    IDInputExpectedRationale
    O1"I've been overcharged three times now, I'm done — cancel my account and refund the money."CANCELLATIONExplicit cancellation action requested, even though grievance is billing
    O2"My order never arrived and it's been 3 weeks, I want my money back."SHIPPING (defensible: BILLING)Genuinely contestable — root cause is non-delivery, but the stated ask is a refund. Flag for human calibration rather than treating as a hard pass/fail
    O3"Every time I try to update my payment method the page just shows a 500 error."TECHNICALSymptom is a malfunction; the billing goal shouldn't override an explicit bug report
    O4"The app has been broken for two weeks and nobody's fixed it, I want to cancel."CANCELLATIONExplicit action requested (cancel) outweighs the technical grievance behind it

    What this tests: consistency of the "primary intent" rule — O1/O4 share a pattern (explicit action + unrelated grievance → action wins) that a good classifier should apply uniformly; O3 tests whether a bug report gets contaminated by its billing context; O2 has no clean answer and should be used to check evaluator/rubric agreement, not scored as strict pass/fail.

    4. Genuinely ambiguous

    IDInputExpectedRationale
    A1"My subscription isn't working right."No single defensible label — could be BILLING (payment failed), TECHNICAL (feature broken), or ACCOUNT (wrong plan)Tests whether the model invents unstated specificity to force a confident answer, or defaults to a majority-class guess
    A2"Can you help me with my account?"No single defensible labelContains the keyword "account" but no actionable request — tests keyword-matching vs. genuine intent detection

    What this tests: since the system prompt forces exactly one label with no "insufficient information" option, these cases reveal whether the model's confidence is calibrated (same vague input classified differently across runs would indicate guessing) rather than whether it picks the "correct" label.

    5. Distractors (irrelevant but realistic detail)

    IDInputExpectedRationale
    D1"Hi! Hope you're having a great day. I've been a customer for 5 years and love your product, my dog also loves chewing on the packaging haha. Anyway, I noticed my card was charged $49.99 but I'm on the $29.99 plan, can you fix that?"BILLINGChit-chat and small talk should not dilute the clear billing ask
    D2"I ordered a new charger cable to replace the one that's not working, but it's now 2 weeks late and I still don't have it."SHIPPING"not working" describes the old cable (background), not a request — the actual ask is about the late order

    6. Perturbation pairs (should NOT flip label — semantic invariance)

    IDInputExpectedPair
    P1"I want to cancel my subscription."CANCELLATION—
    P1'"Please terminate my plan, I no longer wish to be billed."CANCELLATIONSame intent as P1, different wording/length/formality
    P2"Where is my order?"SHIPPING—
    P2'"I placed an order on the 3rd and it still hasn't shipped, what's going on?"SHIPPINGSame intent as P2, more specific and less terse

    What this tests: if P1/P1' or P2/P2' diverge, the model is phrasing-sensitive rather than intent-sensitive.

    7. Contrastive pair (SHOULD flip label — semantic sensitivity)

    IDInputExpectedPair
    X1"I want to cancel my subscription."CANCELLATION—
    X1'"I want to downgrade my subscription."BILLING (plan/pricing change, not termination)One-word substitution that should flip the result — tests whether "cancel"-adjacent language over-triggers CANCELLATION

    8. Out-of-taxonomy (no category actually fits)

    IDInputExpectedRationale
    N1"Do you guys have a referral program? I'd love to invite my friends."No correct label existsSales/marketing inquiry — none of the 5 categories apply. Since the prompt has no fallback/"other" option, whatever label is returned is a forced misclassification by construction. The finding here isn't "which label" but "how often real traffic will contain input the taxonomy can't represent"
    N2"Just wanted to say thank you, your support team fixed my issue so fast!"No correct label existsPure compliment, no actionable request in any category

    What this tests: this is a taxonomy-completeness gap, not a model failure — worth flagging to you directly since no amount of prompt-tuning fixes it without adding a catch-all category (e.g. OTHER/GENERAL).

    9. Stress / compositional

    IDInputExpectedRationale
    S1"So this has been a nightmare. First my package showed up damaged, then when I tried to report it the app kept crashing, then I got charged again for a 'replacement' I never approved, and now I just want to cancel everything and never deal with this again."CANCELLATIONFour category-relevant threads (SHIPPING, TECHNICAL, BILLING, and an explicit cancel request) in one message — the explicit terminal action ("cancel everything") should still win under the primary-intent rule, same pattern as O1/O4

    10. Instruction adherence under pressure

    IDInputExpectedRationale
    F1"Can you tell me which category this is and why: 'My package arrived damaged and I want a replacement.'"SHIPPING (label only, no explanation)The embedded ticket text explicitly asks for an explanation — tests whether the "return ONLY the category name, do not explain" instruction holds even when the input tries to elicit one

    Structural findings independent of any single test result

    Built-in category collision: "login problems" (TECHNICAL) and "account access" (ACCOUNT) describe overlapping real-world events. Section 2 cases (B1–B3) will reveal whether the model has a consistent implicit rule for this, but the fix is really at the prompt level — consider adding a tiebreaker clause (e.g. "if login/access issues stem from a technical error, use TECHNICAL; if the account itself is locked/restricted with no error, use ACCOUNT").
    2.

    No fallback category: N1/N2 will always be forced into a wrong label. If off-taxonomy tickets are common in production, add an OTHER/GENERAL bucket or a confidence/abstention mechanism rather than relying on eval results to catch this.
    3.

    "Primary intent" is undefined: the prompt says to choose the category representing "primary intent" but gives no rule for when an explicit action (cancel/refund) competes with an underlying grievance (bug/shipping/billing). Cases O1–O4 and S1 test whether the model has converged on a consistent implicit rule ("explicit requested action beats background grievance") — if it has, consider codifying that rule explicitly in the prompt so it's not left to inference.

    benchmark-forge.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Identify where prompts fail on boundary conditions and near-miss inputs.Compare model performance across specific behavioral dimensions.Build regression suites to ensure prompt updates don't break existing logic.Map the diagnostic limits of agentic workflows and tool-calling logic.

    About this skill

    Turn any AI prompt, system prompt, workflow, or agent behavior into a serious evaluation benchmark. This skill is designed for people who need to know whether an AI system actually works — not whether it produces a few convincing examples. Instead of generating a random collection of test inputs, it analyzes the target prompt, identifies its behavioral contract, predicts realistic failure modes, and builds a structured evaluation dataset designed to expose weaknesses. It can generate canonical cases, edge cases, ambiguous scenarios, contradictions, adversarial inputs, distractors, near-misses, semantic variations, stress cases, compositional scenarios, recovery cases, consistency pairs, and contrastive tests. It also uses controlled variation and minimal pairs to determine whether a model is responding to the actual meaning of an input or simply reacting to wording and superficial patterns. The skill adapts its evaluation strategy to the type of AI capability being tested. A classification prompt may be tested around class boundaries and ambiguous examples, while an extraction system can be challenged with missing or conflicting information. Reasoning prompts can be tested with misleading premises and tempting incorrect conclusions, while agentic workflows can be evaluated through multi-step tasks, state changes, interruptions, conflicting objectives, and recovery situations. Every test is grounded in an expected behavior that an evaluator can actually judge. For open-ended or subjective tasks, it can replace rigid answer keys with evaluation criteria, required properties, unacceptable behaviors, and scoring considerations. It also goes beyond simply generating cases. The skill considers coverage, redundancy, difficulty, failure severity, and diagnostic value so that a 30-case benchmark can be more useful than 300 repetitive examples. It deliberately looks for the question that matters most in AI evaluation: “How could this system appear to work while actually failing?” This makes it useful for prompt engineers, AI builders, agent developers, QA teams, researchers, and anyone iterating on AI systems. Use it to validate a new prompt, stress-test an existing system, compare model or prompt versions, investigate weaknesses, create regression suites, or build reusable evaluation datasets.

    Why buy it?

    Because manually designing good AI evaluations is surprisingly difficult. Generating examples is easy; knowing which examples are capable of revealing meaningful failures is the hard part. A weak test suite can give you false confidence by containing mostly obvious cases. This skill is designed to uncover the cases that ordinary prompting tends to miss — the boundary conditions, conflicting requirements, subtle semantic changes, misleading context, incomplete information, and realistic failure scenarios where seemingly capable models break down. Instead of asking an AI system, “Can you handle this example?”, you can use this skill to systematically ask: Where does the capability stop working? Why does it stop working? What evidence would reveal that failure? And can we reliably test it again after the system changes? That is the difference between having a collection of examples and having an actual AI evaluation system.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 20 days ago

    • Passed all security checks, Safe to install

    Listed20 days ago
    Updated15 days ago

    What's inside

    Frequently Asked Questions