LLM Eval Framework Builder
You changed the prompt, tried four inputs, it looked better, you shipped — and three days later support tickets say outputs are worse for an entire class of inputs you didn't test
New: Skill bounties are live. Post a request, fund the bounty, and creators compete for 7 days to build it -> See open bounties
THE AGENSI STORE
63 skills found
You changed the prompt, tried four inputs, it looked better, you shipped — and three days later support tickets say outputs are worse for an entire class of inputs you didn't test
by Indy Agent
Evaluate any feature request with structured scoring and a Build / Investigate / Defer / Decline decision.
by TopAgent
Build and curate high-quality prompt datasets for fine-tuning and eval — deduped and labeled.
Evaluate company AI maturity across 6 dimensions with weighted scoring, radar charts, and a GDPR risk audit.
by Timoranjes
Evaluate third-party agent skills for command injection, prompt injection, and data exfiltration before installation.
by GTDataworks
Convert loose prompt sets into structured, target-ready records with variables, contracts, and eval cases.
by Nex AI
Wijst onkostenregels uit een CSV-export automatisch toe aan Belgische MAR-voorbeeldrekeningen op basis van trefwoorden, met review-flags voor twijfelgevallen.
A retrieval architect that diagnoses why RAG returns confident-but-wrong answers, picks the right context architecture (RAG vs knowledge graph vs structured/temporal retrieval) instead of defaulting to vector search, and designs the institutional-memory schema embeddings throw away.
by Kaymue
Diagnose broken RAG systems. 8 failure categories: chunking, embeddings, retrieval, reranking, hallucination. Recall@k measurement.
Design and write the eval suite for your LLM-powered feature — the metrics that match your failure modes, a golden dataset plan with starter cases, anchored rubrics, LLM-as-judge prompts with the known bias mitigations, and pass/fail gates wired for CI.
by Ifásola
Diagnose RAG bottlenecks with precision metrics (Recall, MRR, nDCG) to identify retrieval or ranking failures.
by Echo Rose
Agent Eval Harness - A Premium AI Agent Skill
by Timoranjes
Teaches AI coding agents to build and run automated regression tests for SKILL.md files. When you update a skill that your team depends on, you need to know it still works — not just that it "looks ri
by John Barros
Evaluate and plan the migration of vision inference pipelines to native OpenCV 5 DNN CPU execution.
Model quality is table stakes — the harness is where agents win or fail. This designs yours: it writes a structured, testable system prompt (role, tools, boundaries, method, output contract, failure handling) and maps every concern to the right layer — prompt, tool, guardrail, or evaluation — so the pieces reinforce each other instead of fighting.
by PromptWagon
Designs practical memory architectures for AI assistants, agents, copilots, automations, and workflows, including memory schemas, retention rules, update policies, retrieval keys, summary formats, privacy boundaries, conflict handling, user preference memory, project memory, task memory, and audit-ready memory governance notes for builders.
by PromptWagon
Reviews document sets, source quality, chunking logic, metadata, retrieval coverage, citation traceability, answer grounding, source gaps, stale content, duplicate content, and failure patterns for RAG knowledge-base chatbots. Helps AI, product, support, governance, and engineering teams diagnose common and costly RAG quality problems before deployment or after incidents.
by PromptWagon
Turns model test results, prompts, outputs, benchmarks, scoring notes, evaluation datasets, failure examples, comparison results, and reviewer observations into clear model evaluation reports with findings, recommendations, evidence gaps, deployment considerations, and repeatable evaluation documentation for AI teams.
A deep reference library of production agent patterns — orchestration, context, tool design, failure and recovery, oversight, and evaluation. Every pattern states when it applies, when it's the wrong answer, what it costs, and the failure it prevents. Seven reference files, not a checklist.
Your skill works today. Will it work after the next model update? Build the harness that answers with numbers. Builds a standalone regression test harness with mechanical grading to verify skill behavior after model or code updates.
Replay automation traces against explicit rules for required steps, order, retries, approvals, idempotency, external side effects, duration, and terminal success.
A defensive catalog of ~39 agent security attack patterns across every surface — injection, tool abuse, exfiltration, memory poisoning, multi-agent trust, retrieval poisoning — each with a detection signal, a concrete defense, and a severity. Nine reference files plus a threat-model worksheet. For hardening agents you own.
Audit real RAG evidence traces for missing expected sources, impossible citations, unsupported claims, stale evidence, unused context, and weak routing.
by Shandra
Audits and optimizes checkout and payment experiences for e-commerce stores, SaaS subscriptions, marketplaces, booking platforms, and digital products. Evaluates cart review, pricing transparency, guest checkout, shipping, billing, payment methods, discounts, trust, validation, declines, uncertain payment states, order confirmation, accessibility, analytics, abandoned-checkout recovery, and implementation readiness.