评测对话流程 · Dialogue Evaluator
评测AI对话流程的每一轮:检查回复质量、转折处理与目标达成,输出评分卡与改进清单,适合调Prompt和验收。
$7
Works with the AI tools you already use
评测对话流程 · Dialogue Evaluator
Example session with this skill installed
这是客服机器人和用户的一段8轮对话,最后用户说'算了,我找人工',帮我逐轮评测,看看问题出在哪。
- Read your context and instructions
- Compiled the dialogue evaluator
总评
8 轮对话综合 5.5/10,主要失分在「重复提问」与「目标推进慢」。第 5 轮后用户耐心耗尽,第 7 轮转人工诉求未被识别,属于流程问题。
逐轮评分卡
| 轮次 | 理解 | 推进 | 语气 | 主要问题 |
|---|---|---|---|---|
| 1 | 8 | 8 | 8 | 开场正常 |
| 2 | 7 | 7 | 8 | 未确认订单号就追问 |
| 3 | 4 | 3 | 7 | 与第 2 轮重复提问 |
| 4 | 6 | 5 | 8 | 回答绕,没直接给方案 |
| 5 | 5 | 4 | 6 | 用户已不耐烦,机器人无感 |
| 6 | 4 | 3 | 7 | 继续输出标准话术 |
| 7 | 2 | 1 | 6 | 「找人工」诉求未识别 |
| 8 | 3 | 1 | 7 | 挽留话术激怒用户 |
问题归因
- 流程问题:缺少「情绪检测」与「转人工触发」两个节点,第 5 轮后应直接转人工;
- 提示词问题:第 3 轮重复提问,缺少「已问信息不重复问」的规则;
- 话术问题:第 8 轮挽留不当时,应道歉并立即转接。
改进清单
- 增加情绪关键词检测:出现「算了 / 找人工 / 投诉」立即转接;
- 提示词补充:已收集信息用变量记录,禁止重复提问;
- 转人工前固定话术:「抱歉让您久等,我马上为您转接人工,预计等待 1 分钟。」
下一步
- 先改 3 条规则,拿同一段对话重测;
- 用 20 组真实对话批量评测,统计转人工率;
- 两周后复测,对比评分卡变化。
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
它解决什么问题
搭好 AI 对话流程就上线,常在前几轮跑偏:答非所问、重复提问、卡在死循环、用户暴躁时不会安抚。直接让 AI 自查,会给自己打高分。
这个技能做什么
把整段对话按轮次拆开评测:每轮检查理解是否准确、回复是否推进目标、转折是否平滑;按「目标达成 / 信息完整 / 语气得体 / 效率」打分,输出问题定位与具体改法(改提示词还是改流程)。
为什么比裸提示词强
- 按轮评测:问题定位到具体哪一轮。
- 区分原因:是提示词问题还是流程问题。
- 可复测:同一套评分卡反复用。
典型用例
- 上线前验收客服/销售对话机器人。
- 调提示词后做前后对比评测。
- 把差评对话复盘成改进清单。
局限
需要你提供对话记录;评分卡需按业务目标定制,主观语气项请人工复核。
English · 英文介绍
The problem
Once the AI dialogue process is set up and launched, it often goes astray in the first few rounds: answering questions that are wrong, repeating questions, getting stuck in an endless loop, and not appeasing users when they are irritable. Just let the AI self-check and it will give itself a high score.
What it does
The entire conversation is divided into rounds for evaluation: each round checks whether the understanding is accurate, whether the reply advances the goal, and whether the transition is smooth; score according to "goal achievement/complete information/appropriate tone/efficiency" and output the problem location and specific correction methods (change prompt words or process).
Why is it better than naked prompt words?
- Round-by-round evaluation: The question is located in a specific round.
- Distinguish the reason: is it a prompt word problem or a process problem.
- Retestable: The same set of scorecards can be used repeatedly.
Typical use cases
- Accept customer service/sales dialogue bots before going online.
- After adjusting the prompt words, do a before and after comparison evaluation.
- Review negative review conversations into an improvement list.
Limitations
You are required to provide conversation records; the scorecard needs to be customized according to business goals, and subjective tone items must be reviewed manually.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean today
- Passed all security checks, Safe to install