What Statistical Power Analysis Is: Definition & Uses
- Statistical power analysis calculates the sample size a test needs to reliably detect a real effect; statistical power is the probability of correctly finding a true effect (commonly targeted at 80%).
- It ties together four inputs, effect size, significance level (usually 5%), power (usually 80%), and sample size, and lets you solve for whichever one you don't know (usually sample size).
- Power is the partner to significance: significance guards against false positives (Type I), power guards against false negatives (Type II), missing a real effect.
- The sample size it produces sets the test's duration: required sample ÷ daily eligible traffic per version = days to run; smaller effects and lower traffic mean longer tests.
- An underpowered test both misses real effects AND exaggerates the wins it does catch (the 'winner's curse'), so decide the required sample in advance and run to it.
Most conversations about A/B testing worry about false positives, seeing a win that isn't real. The quieter, more common failure is the opposite: running a test too small to notice a win that is real, then concluding a good idea didn't work. Statistical power analysis is the calculation that prevents it, by telling you, before you launch, how much data the test needs to give a genuine effect a fair chance of showing up. This guide explains what statistical power is, the four inputs a power analysis balances, how it differs from significance, how it sets test duration, and what goes wrong without it, drawing on the experimentation practice behind 70,000+ experiments across 7,000+ websites in 15+ industries [CROBenchmark Report 2026, Omniconvert].
One idea holds it together: an experiment is only worth running if it is big enough to detect the effect you care about, and power analysis is how you make sure it is.
What statistical power analysis is
Statistical power analysis is a calculation that tells you how large a sample an experiment needs in order to reliably detect an effect of a given size, if that effect is really there. Statistical power itself is the probability that a test will correctly find a true effect, in other words, that it will not miss a real difference. A commonly used target is 80% power: if a genuine effect of the size you care about exists, the test has an 80% chance of detecting it as statistically significant, and a corresponding 20% chance of missing it.
Power analysis ties together four quantities and lets you solve for whichever one you don't know. In practice, its most important use is to work out the sample size, and therefore the duration, a test needs before you run it. To see how, look at the four quantities it balances.
The four inputs to a power analysis
A power analysis connects four quantities, and if you fix any three, it gives you the fourth:
- Effect size. The smallest difference worth detecting, the minimum detectable effect (e.g. a lift from a 4% to a 4.4% conversion rate). Smaller effects are harder to detect and need bigger samples.
- Significance level (alpha). Your tolerance for a false positive, conventionally 5%. A stricter level makes the test more demanding.
- Power. Often 80% or 90%, your tolerance for a false negative, the probability of catching a real effect if it exists.
- Sample size. How many observations (visitors, sessions) each group needs, usually the quantity you solve for.
The usual workflow is to decide the effect size, significance level, and power, then solve for the sample size, which tells you how much traffic the test needs and therefore how long it must run. The two inputs people most often confuse, significance and power, are worth separating clearly.
Power vs statistical significance
| Dimension | Statistical significance | Statistical power |
|---|---|---|
| Error it guards against | False positive (Type I): seeing an effect that isn't there | False negative (Type II): missing an effect that is |
| Typical target | 5% significance level (95% confidence) | 80% power |
| When it's handled | Judged after the test, from the data (p-value) | Planned before the test, to size the sample |
| Question it answers | Is this difference unlikely to be chance? | Is the test big enough to find a real difference? |
Because power is planned in advance, its most practical output is the one thing every test manager needs: how long to run.
How power analysis sets test duration
Power analysis sets test duration indirectly, by first determining the required sample size and then translating that into time using your traffic. Decide the minimum effect you care about detecting, choose your significance level (typically 5%) and desired power (typically 80%), and run the power analysis to get the sample size each group needs. Then divide that required sample by how many eligible visitors the tested page sends to each version per day to estimate how many days the test must run.
For example, if the analysis says each variation needs 25,000 visitors and the page sends 2,500 eligible visitors a day to each version, the test needs roughly ten days, and you would round up to run across whole weeks to capture day-of-week variation. This is why a smaller expected effect, or a lower-traffic page, means a longer test duration. And it explains what goes wrong when this step is skipped.
What happens if a test is underpowered
An underpowered test, one whose sample is too small for the effect you're trying to detect, has a high chance of missing a real effect, and it causes damage in two directions. Most obviously, it produces false negatives: a genuinely better variation fails to reach significance, you record it as "no difference," and a valuable improvement is discarded because the experiment never had the statistical muscle to reveal it. Over time, a program of underpowered tests quietly throws away good ideas and makes it look as though "nothing works."
Less obviously, underpowered tests distort even their apparent wins: when a small test does happen to cross the significance line, it tends to do so only on an unusually large observed effect, so the estimated size of the improvement is exaggerated, a phenomenon sometimes called the "winner's curse." The fix is the same discipline power analysis exists to enforce, and it is exactly what a good platform makes routine.
Adequately powered tests with Omniconvert Explore
Omniconvert Explore is an A/B testing and experimentation platform built to keep tests statistically honest, and running adequately powered experiments is central to that. The principle it enforces is the one power analysis exists to serve: decide how much data a test needs before you launch, then run to that sample rather than stopping the moment a result looks good. Explore reports statistical significance and confidence as data accumulates, so you can see when a test has genuinely reached a conclusive result instead of reacting to an early, underpowered blip.
Its segmentation is also relevant to power in a practical way: a test that is adequately powered for your whole audience is usually underpowered for any single narrow segment, because each segment has a fraction of the traffic, so Explore's segment-level view helps you judge which segment findings have enough data behind them to trust and which are too thin to act on. Across more than 70,000 experiments, with an average uplift of 23.2%, this discipline, size the test properly and run it to completion, is what turns experimentation into a reliable source of real, repeatable wins.
Give every real improvement a fair chance to prove itself.
See how Omniconvert Explore runs tests to a conclusive result →Frequently Asked Questions
Statistical power analysis is a calculation that tells you how large a sample an experiment needs in order to reliably detect an effect of a given size, if that effect is really there. Statistical power itself is the probability that a test will correctly find a true effect, that it will not miss a real difference. A commonly used target is 80% power, meaning that if a genuine effect of the size you care about exists, the test has an 80% chance of detecting it as statistically significant (and a 20% chance of missing it, a 'false negative'). Power analysis ties together four quantities, effect size, significance level, power, and sample size, and lets you solve for whichever one you don't know. Its most important use is to work out the sample size (and therefore duration) a test needs before you run it. Running a test without one is like sending out a survey without checking whether you asked enough people: even a real effect can slip through undetected because the sample was too small.
A power analysis connects four quantities, and if you fix any three, it gives you the fourth. The first is the effect size, the smallest difference you care about detecting (the minimum detectable effect, e.g. a lift from a 4% to a 4.4% conversion rate); smaller effects need bigger samples. The second is the significance level (alpha), your tolerance for a false positive, conventionally 5%. The third is the power itself (often 80% or 90%), your tolerance for a false negative, the probability of catching a real effect. The fourth is the sample size, how many observations each group needs. Because these four are mathematically linked, the usual workflow is to decide the effect size, significance level, and power you want, then solve for the sample size, which tells you how much traffic the test needs and how long it must run. The key intuition: the smaller the effect you want to catch and the more certain you want to be of catching it, the more data you need.
Because an underpowered A/B test is close to useless: it can fail to detect a real, valuable improvement simply because it did not collect enough data, and then you wrongly conclude a good change 'didn't work.' Most people worry about false positives, calling a difference real when it is only noise, but power addresses the opposite and equally costly error: false negatives, missing a genuine effect. If a test has only 50% power for the effect you care about, then even when that effect is real you have a coin-flip's chance of missing it, so half your good ideas could be discarded on the basis of an experiment that never had a fair chance. Running a power analysis before the test tells you the sample size (and duration) needed to give a real effect a strong chance of showing up, typically 80% or more. It also protects against stopping a test as soon as it looks significant, which breaks the logic power is calculated on and inflates false positives.
They guard against two different errors, and a good experiment needs both. Statistical significance concerns false positives (Type I errors): it asks whether an observed difference is unlikely to have arisen by chance, and the significance level (usually 5%) caps how often you call a difference real when it is not. Statistical power concerns false negatives (Type II errors): it asks whether the test is capable of detecting a real difference, and measures the probability of catching a genuine effect if one exists (commonly 80%). Significance protects you from being fooled by noise into seeing an effect that isn't there; power protects you from missing an effect that is. You can have one without the other, a test can use a strict 5% significance level and still be badly underpowered. Significance is judged after the test from the data (via the p-value and the null hypothesis); power is planned before the test to size it correctly.
Indirectly, by first determining the required sample size and then translating that into time using your traffic. The workflow: decide the minimum effect you care about detecting, choose your significance level (typically 5%) and desired power (typically 80%), and run the power analysis to get the sample size each group needs. Then divide that required sample by how many eligible visitors the tested page sends to each version per day to estimate how many days the test must run. For example, if the analysis says each variation needs 25,000 visitors and the page sends 2,500 eligible visitors a day to each version, the test needs roughly ten days, rounded up to run across whole weeks to capture day-of-week variation. This is why a smaller expected effect, or a lower-traffic page, means a longer test: detecting a subtle difference reliably requires more data, which takes more time to accumulate. You fix this duration before launching and run to it, rather than stopping the moment the result looks good.
An underpowered test, one whose sample is too small for the effect you're trying to detect, has a high chance of missing a real effect, and it causes damage in two directions. Most obviously, it produces false negatives: a genuinely better variation fails to reach significance, you record it as 'no difference,' and a valuable improvement is discarded because the experiment never had the statistical muscle to reveal it. Over time, a program of underpowered tests quietly throws away good ideas and makes it look as though 'nothing works.' Less obviously, underpowered tests exaggerate even their apparent wins: when a small test does cross the significance line, it tends to do so only on an unusually large observed effect, so the estimated size of the improvement is overstated (the 'winner's curse'). So they both miss real effects and overstate the ones they catch. The fix is to run a power analysis before the test, size it for the smallest effect worth detecting, and run to that sample.
Omniconvert Explore is an A/B testing and experimentation platform built to keep tests statistically honest, and running adequately powered experiments is central to that. The principle it enforces is the one power analysis exists to serve: decide how much data a test needs before you launch, then run to that sample rather than stopping the moment a result looks good. Explore reports statistical significance and confidence as data accumulates, so you can see when a test has genuinely reached a conclusive result instead of reacting to an early, underpowered blip, precisely the peeking-and-stopping trap that undermines power. Its segmentation is also relevant: a test adequately powered for your whole audience is usually underpowered for any single narrow segment, because each segment has a fraction of the traffic, so the segment-level view helps you judge which segment findings have enough data to trust. Across more than 70,000 experiments, with an average uplift of 23.2%, this discipline, size the test properly and run it to completion, turns experimentation into a reliable source of real, repeatable wins.
Statistical power analysis answers the question you should ask before every experiment: is this test big enough to find what I'm looking for? Power is the probability of correctly detecting a real effect (commonly targeted at 80%), and a power analysis ties together four quantities, effect size, significance level, power, and sample size, letting you solve for the one you need, almost always the sample size, and therefore the duration. It is the essential partner to statistical significance: significance guards against false positives (seeing an effect that isn't there), power guards against false negatives (missing an effect that is). Skip it and you run underpowered tests that quietly discard good ideas, make it look as though 'nothing works,' and even exaggerate the wins they do catch. Do it, and you give every real improvement a fair chance to prove itself. The discipline is simple, decide the required sample size in advance and run to it rather than stopping at the first good-looking moment, and it is exactly what Omniconvert Explore is built to enforce across every experiment.
Run tests that can actually find a winner with Omniconvert Explore
An underpowered test throws away good ideas. Omniconvert Explore reports significance and confidence as data accumulates, so you run experiments to a conclusive result instead of stopping early, and segment to see which findings have the data to back them.