ER

    评测报告编排 · Eval Report Orchestrator

    1

    把模型评测结果编排成结论先行的报告:指标对比、失败案例与改进优先级。 Turn model evaluation results into a conclusion-first report with metric comparisons, failure cases, and improvement priorities.

    $7

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    0 installsSecurity scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more
    ER

    评测报告编排 · Eval Report Orchestrator

    Example session with this skill installed

    请把 300 条测试集的评测结果整理成报告,对比 v2/v3 提示词,给出改进优先级。

    • Read your context and instructions
    • Compiled the eval report orchestrator

    评测报告编排完成。

    结论

    v3 在准确率上领先 v2 3.2 个百分点,但在长文本任务上出现 12 例超时,暂不建议全量。

    指标对比

    指标v2v3差异
    准确率84.1%87.3%+3.2pp
    事实一致91.0%92.6%+1.6pp
    超时率0.8%4.0%-3.2pp

    失败案例

    • 长文档摘要:12 例超时,集中在 8k 以上输入
    • 多跳问答:4 例漏引证据

    改进优先级

    1. 超时:拆分长文本调用(高影响)
    2. 多跳问答:补 20 条针对性样本(中)
    3. 其余指标维持

    Next steps

    1. 超时修复后重跑长文本子集。
    2. v3 与修复版做单盲对比。

    Connects securely to your tools. The creator never sees your data.

    What you get

    Structure raw eval tables into executive summariesIdentify failure patterns in LLM prompt versionsPrioritize model improvements based on impact scoresGenerate weekly regression reports for dev teams

    About this skill

    它解决什么问题

    评估结果以原始评分表形式出现,没有任何叙述,因此团队无法判断哪个模型或即时更改真正重要。

    这个技能做什么

    • 将评估结果构建为结论优先的报告。
    • 比较相同指标的模型和提示版本。
    • 提取故障案例和修复模式。
    • 按预期影响对改进进行排名。

    为什么比裸提示词强

    它强制执行评估报告框架:结论、指标、失败、原因、优先级,因此审查以排名的行动列表结束。

    典型用例

    • 每周法学硕士回归报告。
    • 及时进行 A/B 版本评估。
    • 预发布模型验证审查。

    局限

    它组织您的评估数据;它不运行评估。


    English · 英文介绍

    The problem

    Eval results arrive as raw score tables with no narrative, so teams cannot tell which model or prompt change actually matters.

    What it does

    • Structures eval results into a conclusion-first report.
    • Compares models and prompt versions on the same metrics.
    • Extracts failure cases and patterns for fixes.
    • Ranks improvements by expected impact.

    Why this beats prompting it yourself

    It enforces an eval report frame: verdict, metrics, failures, causes, priorities, so reviews end with a ranked action list.

    Use cases

    • Weekly LLM regression reports.
    • Prompt version A/B evaluations.
    • Pre-release model validation reviews.

    Known limitations

    It organizes your eval data; it does not run the evaluations.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    • Passed all security checks, Safe to install

    Listedtoday

    Frequently Asked Questions