Definition
Multiple testing correction adjusts decision thresholds or p-values for a defined family of statistical tests. Family-wise error methods address the chance of at least one false rejection, while false discovery rate methods address the expected proportion of false discoveries among rejections. The method must match the intended error criterion and its assumptions.
Why It Matters
- A merchandising test can generate dozens of comparisons across categories and metrics. Selecting only the smallest p-value hides how many opportunities there were to find a result.
- The Commerce Intelligence OS framing calls for a traceable promotion decision: which evidence justified keeping an intervention?
- Teams need a record of exploratory findings so a later confirmatory test can target a specific commercial question.
How It Works
- Before examining outcomes, define the primary question and the family of comparisons. Record which segment analyses are exploratory.
- Choose an appropriate procedure: Bonferroni divides the family error level by the number of tests; Holm and false discovery rate methods offer other controls with different properties.
- Report the full comparison family, effect estimates, and adjusted results. Check the chosen procedure against the test dependence structure.
- Keep the rollout decision attached to the original experiment brief. Evaluate practical benefit and mature return outcomes; a statistical adjustment does not repair biased assignment or repeated unplanned peeking.
Ecommerce Example
Context: Illustrative example: a retailer predefines 20 category comparisons and chooses Bonferroni control at a family level of 0.05.
Recommended move: The per-comparison threshold is 0.05 divided by 20, or 0.0025. An unadjusted p-value of 0.03 does not meet that rule.
Why it matters: The category can remain a documented follow-up hypothesis. These hypothetical figures do not establish any commercial uplift.
iKawn Framework
Register
The iKawn framework records questions and comparison families before evaluation.
Evaluate
Retain every tested comparison, including inconclusive results.
Qualify
Attach the chosen error criterion to the evidence presented to decision-makers.
Confirm
Convert promising exploratory findings into a specific subsequent test.
Concise Summary
Many comparisons need a declared decision rule. Preserve the full test family and separate exploratory signals from evidence used to authorize rollout.