Definition
Equivalence testing asks whether an effect lies within prespecified acceptable bounds. The two one-sided tests approach tests both boundaries. Failing to detect a difference in an ordinary superiority test does not by itself establish equivalence.
Why It Matters
- A merchant replacing an expensive workflow may need evidence that a key customer outcome stays within an acceptable range.
- A Commerce Intelligence OS should preserve the business tolerance behind a decision rather than label an inconclusive experiment a tie.
How It Works
- Define the effect measure and acceptable lower and upper bounds before reading outcomes. Explain why those bounds are commercially meaningful.
- Plan sample size and assignment around the equivalence question. Use analysis that respects the actual experimental unit and dependence structure.
- Test both boundaries using an appropriate method. With the standard TOST setup at 5% per one-sided test, the matching 90% confidence interval must fall wholly inside the bounds.
- Report the interval, bounds, and other decision guardrails. An inconclusive result should remain inconclusive rather than authorize rollout by default.
Ecommerce Example
Context: Illustrative example: a retailer defines acceptable change in a customer rating score as between -0.2 and +0.2 points before testing two support flows.
Recommended move: A compatible 90% interval from -0.1 to +0.15 lies inside those bounds; an interval from -0.3 to +0.1 does not.
Why it matters: This illustrates the decision rule under a suitable design. Cost savings, return outcomes, and other guardrails still require separate evidence.
iKawn Framework
Specify
The iKawn framework connects tolerance bounds to a commercial decision.
Design
Record the unit of assignment and planned analysis.
Assess
Compare uncertainty with the approved bounds.
Decide
Keep equivalence evidence distinct from superiority and operational guardrails.
Concise Summary
Equivalence needs predefined bounds and enough precision to rule out unacceptable differences. A nonsignificant superiority test is insufficient.