Skip to content

Sample Size and Statistical Power

How to plan experiments to get reliable results.

What is Sample Size

Sample size is the number of users (or other randomization units) that need to be included in the experiment. A sample that is too small will not give statistically significant results, even if there is an effect.

The correct sample size depends on four parameters: baseline metric value, expected change, statistical power, and significance level.

Statistical Power

Statistical power (power) is the probability of detecting an effect if it actually exists. Usually set at 80%.

If power is 80%, this means: when there is a real effect, you will detect it in 80% of cases. In the remaining 20% of cases, the result may not be significant due to randomness (false negative result).

Expected Change (MDE)

Minimum Detectable Effect (MDE) is the minimum metric change you want to detect. The smaller the expected change, the larger the sample needed.

For example, detecting a 1% change in conversion is harder (requires more users) than detecting a 10% change. Therefore, it's important to realistically assess what effect you expect.

Significance Level (Alpha)

Significance level (alpha) is the acceptable probability of false positive (false positive result). Usually set at 5% (0.05).

This means: if there is actually no effect, in 5% of cases the test will erroneously show that there is one. Lowering alpha (for example, to 1%) requires a larger sample.

AB-Labz - Product Experiments Laboratory