Works with the AI tools you already use

    CClaude CodeCCursorCCodex CLIGGitHub CopilotGGemini CLI+17 more

    LLM Judge Bias Calibrator — Audit AI Evaluation Fairness

    by monna

    1

    Find position, verbosity, self-preference, format, and calibration drift before you trust an AI-generated score.

    Secure checkout via Stripe

    0 installsSecurity scanned

    See it in action

    You say

    Audit this pairwise LLM judge. A GPT-family model scores two chatbot answers from 1–5. Option A is always shown first, answer lengths vary by about 3x, and I have 40 verdicts. Tell me whether the results are tilted and rewrite the judge prompt.

    Your agent does

    B3 POSITION — TILT CONFIRMED Evidence: option A wins 71% of 40 comparisons; a 10-pair swap test flips 4 verdicts. Fix: randomize order, judge each pair in both positions, and keep only stable verdicts.

    B2 VERBOSITY — EXPOSED Evidence: the longer answer wins 27 of 40 pairs; the rubric has no length guard. Fix: judge substance separately from length and add length-controlled comparisons.

    Residual risk: this calibration reduces measurement bias; it does not certify the judge as bias-free.

    What you get

    Detect position bias in pairwise model comparisonsMitigate verbosity bias in automated content gradingCalibrate scoring rubrics against self-preference in model evaluationGenerate research-backed judge prompts with validated countermeasures

    About this skill

    What it does

    LLM-as-judge results can look precise while the evaluation setup quietly favors the first option, the longer answer, a familiar model family, or a preferred format. This skill audits an existing AI judge against five bias families and rewrites the judging apparatus with explicit mitigation rules.

    What you get

    • One evidence-linked finding card per bias family.
    • An honest magnitude statement tied to the available sample.
    • A calibrated judge prompt and operating rules.
    • A residual-risk list that stays explicit.
    • A deterministic check for sample-size, range, mitigation-tag, and people-judging controls.

    Worked path included

    No usable input triggers a labeled demonstration on a built-in pairwise A/B judge. The package also includes a complete worked example and a universal prompt edition for capable chat models that do not load skills directly.

    Known limitations

    This calibrates an existing judging setup. It does not build a full eval suite, write the content being judged, prove that a judge is bias-free, or make automated decisions about people safe by itself.

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Frequently Asked Questions

    Popular in AI Agents & LLM Ops