- Home
- Skills
- Research & Analysis
- 校验A/B实验结果 · A/B Test Validator
校验A/B实验结果 · A/B Test Validator
校验A/B实验的数据质量与结论:样本量、显著性、辛普森悖论与偷看风险逐项检查,给出可信度评级与复算清单。 Validate A/B test data quality and conclusions: sample size, significance, Simpson's paradox, and peeking risks, with a confidence rating and recheck list.
$7
Works with the AI tools you already use
校验A/B实验结果 · A/B Test Validator
Example session with this skill installed
我们的A/B实验报告说新版注册按钮让转化率从10%提升到11.2%,实验跑了两天,总样本2000人,帮我校验这个结论可不可信。
- Read your context and instructions
- Compiled the a b a
校验结论
该实验结论目前「暂不可上线」:提升幅度未达到统计显著,且样本量不足以支持 1.2 个百分点的差异判断。建议延长实验或扩大样本后再决策。
一、设计核查
| 检查项 | 结果 | 说明 |
|---|---|---|
| 分组随机化 | 通过 | 按用户 ID 均分 |
| 实验时长 | 风险 | 仅 2 天,未覆盖完整周周期 |
| 提前偷看 | 风险 | 中途多次查看会抬高误判概率 |
二、统计复算
- 对照组:1000 人,转化 100 人,转化率 10.0%;
- 实验组:1000 人,转化 112 人,转化率 11.2%;
- 两比例 z 检验:p ≈ 0.38(未达 0.05 显著水平);
- 置信区间:提升幅度 95% CI ≈ -1.5% ~ +3.9%,区间包含 0。
三、结论判定
- 判定:数据不足,不能宣布有效;
- 可信度评级:C(不可上线);
- 若坚持上线,决策依据应从「实验显著」改为「成本低 + 无负面指标」。
四、复算清单
- 用实验平台导出完整原始数据,核对分流是否均衡;
- 计算最小样本量:要检测 1.2 个百分点提升,约需每组 1.2 万人;
- 延长实验至至少一个完整周周期,期间不查看结果。
下一步
- 把对照组与实验组的分流日志发我,做均衡性检验;
- 确认实验期间是否有营销活动叠加干扰;
- 按「每组 1.2 万人」重新排期实验。
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
它解决什么问题
团队看到实验「显著」就急着上线,常忽略样本量不足、提前偷看结果、分组不随机等问题;也有人在数据里挑子群体,把噪声当成发现。直接问 AI,常得到「显著就有效」的草率结论。
这个技能做什么
把实验数据与设计描述输入后,按「设计核查 → 数据质量 → 统计复算 → 结论判定」四步校验:检查随机化、样本量、显著性水平、置信区间与多指标偷看风险,输出可信度评级与需要人工确认的清单。
为什么比裸提示词强
- 防误判:强制检查显著性、样本量与偷看风险。
- 可复算:给出核心指标的计算口径与公式。
- 结论分级:区分「可上线」「需继续实验」「数据无效」。
典型用例
- 上线前复核 A/B 实验结论是否成立。
- 审计历史实验数据的统计口径。
- 为增长团队建立实验验收标准。
局限
需要你提供原始数据或汇总指标;无法自动抓取实验平台数据,复杂分层模型建议由数据同事复核。
English · 英文介绍
The problem
When the team sees that the experiment is "significant", they rush to launch it. They often ignore problems such as insufficient sample size, peeking at the results in advance, and non-random grouping. Some people also select subgroups in the data and treat noise as discovery. When asking AI directly, we often get the hasty conclusion that "it is effective if it is obvious".
What it does
After entering the experimental data and design description, press "Design Verification → Data Quality → Statistical Recalculation → Conclusion Judgment" to perform four-step verification: check randomization, sample size, significance level, confidence interval and multi-index peek risk, and output a credibility rating and a list that requires manual confirmation.
Why is it better than naked prompt words?
- Prevent misjudgment: Mandatory checking of significance, sample size and peek risk.
- Recalculable: The calculation caliber and formula of the core indicators are given.
- Conclusion classification: distinguish between "can go online", "need to continue experimentation" and "invalid data".
Typical use cases
- Review whether the A/B experiment conclusions are valid before going online.
- Audit the statistical caliber of historical experimental data.
- Establish experiment acceptance criteria for growth teams.
Limitations
You are required to provide original data or summary indicators; experimental platform data cannot be automatically captured, and complex hierarchical models are recommended to be reviewed by data colleagues.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean today
- Passed all security checks, Safe to install