Works with the AI tools you already use
LLM Judge Bias Calibrator — Audit AI Evaluation Fairness
by monna
Find position, verbosity, self-preference, format, and calibration drift before you trust an AI-generated score.
Secure checkout via Stripe
See it in action
You say
Audit this pairwise LLM judge. A GPT-family model scores two chatbot answers from 1–5. Option A is always shown first, answer lengths vary by about 3x, and I have 40 verdicts. Tell me whether the results are tilted and rewrite the judge prompt.
Your agent does
B3 POSITION — TILT CONFIRMED Evidence: option A wins 71% of 40 comparisons; a 10-pair swap test flips 4 verdicts. Fix: randomize order, judge each pair in both positions, and keep only stable verdicts.
B2 VERBOSITY — EXPOSED Evidence: the longer answer wins 27 of 40 pairs; the rubric has no length guard. Fix: judge substance separately from length and add length-controlled comparisons.
Residual risk: this calibration reduces measurement bias; it does not certify the judge as bias-free.
What you get
About this skill
What it does
LLM-as-judge results can look precise while the evaluation setup quietly favors the first option, the longer answer, a familiar model family, or a preferred format. This skill audits an existing AI judge against five bias families and rewrites the judging apparatus with explicit mitigation rules.
What you get
- One evidence-linked finding card per bias family.
- An honest magnitude statement tied to the available sample.
- A calibrated judge prompt and operating rules.
- A residual-risk list that stays explicit.
- A deterministic check for sample-size, range, mitigation-tag, and people-judging controls.
Worked path included
No usable input triggers a labeled demonstration on a built-in pairwise A/B judge. The package also includes a complete worked example and a universal prompt edition for capable chat models that do not load skills directly.
Known limitations
This calibrates an existing judging setup. It does not build a full eval suite, write the content being judged, prove that a judge is bias-free, or make automated decisions about people safe by itself.
How to install
Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean today
- 30-day refund guarantee
- One-time purchase, yours forever
- Secure checkout via Stripe
Frequently Asked Questions
Popular in AI Agents & LLM Ops
designing-hybrid-context-layers
Architects the right retrieval strategy for every query — teaching your agent when to use RAG, a knowledge graph, or a temporal index instead of defaulting to vector search for everything.

AI Coding Prompt Refiner for Better Developer Results
Transforms vague coding requests into precise, scoped, testable, AI-ready prompts for Cursor, Claude Code, Codex CLI, Replit, and other coding agents.

prompt-engineer
Professional prompt engineering patterns for building robust, secure, and production-ready LLM applications.
skill-creator
The "Skill for building Skills": Automate creating, testing, and optimisation of custom workflows.