2.7w phuryn

ab-test-analysis Skill

分析 A/B 测试结果,包含统计显著性、样本量校验、置信区间,以及上线/延长/停止的建议。适用于评估实验结果、判断测试是否达到显著性、解读分流测试数据,或决定是否上线某个变体。

安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。

查看源码

技能指令原文(SKILL.md)

A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

Context

You are analyzing A/B test results for $ARGUMENTS.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

Instructions

  1. Understand the experiment:
  • What was the hypothesis?
  • What was changed (the variant)?
  • What is the primary metric? Any guardrail metrics?
  • How long did the test run?
  • What is the traffic split?
  1. Validate the test setup:
  • Sample size: Is the sample large enough for the expected effect size?
  • Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
  • Flag if the test is underpowered (<80% power)
  • Duration: Did the test run for at least 1-2 full business cycles?
  • Randomization: Any evidence of sample ratio mismatch (SRM)?
  • Novelty/primacy effects: Was there enough time to wash out initial behavior changes?
  1. Calculate statistical significance:
  • Conversion rate for control and variant
  • Relative lift: (variant - control) / control × 100
  • p-value: Using a two-tailed z-test or chi-squared test
  • Confidence interval: 95% CI for the difference
  • Statistical significance: Is p < 0.05?
  • Practical significance: Is the lift meaningful for the business?

If the user provides raw data, generate and run a Python script to calculate these.

  1. Check guardrail metrics:
  • Did any guardrail metrics (revenue, engagement, page load time) degrade?
  • A winning primary metric with degraded guardrails may not be a true win
  1. Interpret results:

| Outcome | Recommendation |
|---|---|
| Significant positive lift, no guardrail issues | Ship it — roll out to 100% |
| Significant positive lift, guardrail concerns | Investigate — understand trade-offs before shipping |
| Not significant, positive trend | Extend the test — need more data or larger effect |
| Not significant, flat | Stop the test — no meaningful difference detected |
| Significant negative lift | Don't ship — revert to control, analyze why |

  1. Provide the analysis summary:
   ## A/B Test Results: [Test Name]

   **Hypothesis**: [What we expected]
   **Duration**: [X days] | **Sample**: [N control / M variant]

   | Metric | Control | Variant | Lift | p-value | Significant? |
   |---|---|---|---|---|---|
   | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
   | [Guardrail] | ... | ... | ... | ... | ... |

   **Recommendation**: [Ship / Extend / Stop / Investigate]
   **Reasoning**: [Why]
   **Next steps**: [What to do]

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.


Further Reading