DE

    评测对话流程 · Dialogue Evaluator

    0

    评测AI对话流程的每一轮:检查回复质量、转折处理与目标达成,输出评分卡与改进清单,适合调Prompt和验收。

    $7

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    0 installsSecurity scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more
    DE

    评测对话流程 · Dialogue Evaluator

    Example session with this skill installed

    这是客服机器人和用户的一段8轮对话,最后用户说'算了,我找人工',帮我逐轮评测,看看问题出在哪。

    • Read your context and instructions
    • Compiled the dialogue evaluator

    总评
    8 轮对话综合 5.5/10,主要失分在「重复提问」与「目标推进慢」。第 5 轮后用户耐心耗尽,第 7 轮转人工诉求未被识别,属于流程问题。

    逐轮评分卡

    轮次理解推进语气主要问题
    1888开场正常
    2778未确认订单号就追问
    3437与第 2 轮重复提问
    4658回答绕,没直接给方案
    5546用户已不耐烦,机器人无感
    6437继续输出标准话术
    7216「找人工」诉求未识别
    8317挽留话术激怒用户

    问题归因

    • 流程问题:缺少「情绪检测」与「转人工触发」两个节点,第 5 轮后应直接转人工;
    • 提示词问题:第 3 轮重复提问,缺少「已问信息不重复问」的规则;
    • 话术问题:第 8 轮挽留不当时,应道歉并立即转接。

    改进清单

    1. 增加情绪关键词检测:出现「算了 / 找人工 / 投诉」立即转接;
    2. 提示词补充:已收集信息用变量记录,禁止重复提问;
    3. 转人工前固定话术:「抱歉让您久等,我马上为您转接人工,预计等待 1 分钟。」

    下一步

    1. 先改 3 条规则,拿同一段对话重测;
    2. 用 20 组真实对话批量评测,统计转人工率;
    3. 两周后复测,对比评分卡变化。

    Connects securely to your tools. The creator never sees your data.

    What you get

    Audit customer service or sales bots before deployment.Compare prompt performance before and after iterations.Transform negative user feedback into a prioritized fix list.Identify logic loops and dead ends in complex chat flows.

    About this skill

    它解决什么问题

    搭好 AI 对话流程就上线,常在前几轮跑偏:答非所问、重复提问、卡在死循环、用户暴躁时不会安抚。直接让 AI 自查,会给自己打高分。

    这个技能做什么

    把整段对话按轮次拆开评测:每轮检查理解是否准确、回复是否推进目标、转折是否平滑;按「目标达成 / 信息完整 / 语气得体 / 效率」打分,输出问题定位与具体改法(改提示词还是改流程)。

    为什么比裸提示词强

    • 按轮评测:问题定位到具体哪一轮。
    • 区分原因:是提示词问题还是流程问题。
    • 可复测:同一套评分卡反复用。

    典型用例

    • 上线前验收客服/销售对话机器人。
    • 调提示词后做前后对比评测。
    • 把差评对话复盘成改进清单。

    局限

    需要你提供对话记录;评分卡需按业务目标定制,主观语气项请人工复核。


    English · 英文介绍

    The problem

    Once the AI ​​dialogue process is set up and launched, it often goes astray in the first few rounds: answering questions that are wrong, repeating questions, getting stuck in an endless loop, and not appeasing users when they are irritable. Just let the AI ​​self-check and it will give itself a high score.

    What it does

    The entire conversation is divided into rounds for evaluation: each round checks whether the understanding is accurate, whether the reply advances the goal, and whether the transition is smooth; score according to "goal achievement/complete information/appropriate tone/efficiency" and output the problem location and specific correction methods (change prompt words or process).

    Why is it better than naked prompt words?

    • Round-by-round evaluation: The question is located in a specific round.
    • Distinguish the reason: is it a prompt word problem or a process problem.
    • Retestable: The same set of scorecards can be used repeatedly.

    Typical use cases

    • Accept customer service/sales dialogue bots before going online.
    • After adjusting the prompt words, do a before and after comparison evaluation.
    • Review negative review conversations into an improvement list.

    Limitations

    You are required to provide conversation records; the scorecard needs to be customized according to business goals, and subjective tone items must be reviewed manually.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    • Passed all security checks, Safe to install

    Listedtoday

    Frequently Asked Questions