What Sample Size Is: Definition, Why It Matters & How to Set It
- Sample size is the number of observations (n) in a sample; it governs how much random chance affects your result and therefore how much you can trust it.
- A larger sample reduces chance variation and narrows the margin of error, but beyond a point the extra precision isn't worth the extra cost.
- The size you need depends on desired precision and confidence, population variability, and, in experiments, the minimum detectable effect and statistical power, standard inputs a calculator turns into a target.
- Sample size controls precision, not representativeness: a large sample drawn badly is still biased, so a good sample must be both unbiased and large enough.
- In A/B testing, decide the sample size up front and reach it before concluding, calling a test early on a noisy lead is a classic, costly mistake.
Almost every number you report comes from a sample, and how many observations that sample holds quietly decides how much the number can be believed. Sample size is that count. Too few, and random chance runs the show; enough, and the picture steadies into something you can act on. It is one of the least glamorous decisions in any study and one of the most consequential. This guide explains what sample size is, why it matters, what determines the size you need, the problems at both extremes, and its role in A/B testing, and how Omniconvert Explore supports the discipline it demands, drawing on 70,000+ experiments across 7,000+ websites in 15+ industries [CROBenchmark Report 2026, Omniconvert].
One idea holds it together: sample size buys precision, not representativeness, and the right size is the smallest one that answers your question reliably.
What sample size is
Sample size is the number of observations, people, sessions, orders, or other units, included in a sample, usually written as n. When you cannot study a whole population, you study a sample of it, and the sample size is simply how many members of the population are in that sample. It is one of the most important decisions in any study, survey, or experiment, because it governs how much you can trust the result.
A larger sample size reduces the effect of random chance, so estimates from a big sample vary less and sit closer to the true population value; a smaller sample is noisier and can swing widely by luck alone. But bigger is not automatically better, larger samples cost more time and money, and beyond a certain point the extra precision is not worth the extra cost. Importantly, sample size controls precision, not representativeness: a big sample drawn badly is still biased, so a good sample must be both unbiased in how it was selected and large enough to be precise. That precision is exactly why it matters.
Why sample size matters
Sample size matters because it determines how reliable and how precise your results are, and therefore whether you can act on them safely. With a small sample, random chance has a big influence: the result you see could easily be a fluke that would not repeat, so conclusions drawn from it are shaky. As the sample grows, chance variation shrinks, the margin of error around your estimate narrows, and the result becomes something you can trust.
In experiments such as A/B tests, sample size is what gives a test the statistical power to detect a real difference when one exists; too small a sample can miss a genuine improvement (a false negative) simply because there was not enough data to see it. There is a cost to getting it wrong in either direction: too small, and you either miss real effects or chase false ones; needlessly large, and you waste time and money. Choosing the right sample size up front is what lets a study reach a conclusion that is both trustworthy and efficient, which raises the question of how you choose it.
What determines the sample size you need
Several factors together determine the sample size you need, and they trade off against one another:
| Factor | Effect on required sample size |
|---|---|
| Desired precision (margin of error) | Smaller margin of error → larger sample |
| Confidence level | Higher confidence (e.g. 95% vs 90%) → larger sample |
| Population variability | More varied / diverse data → larger sample |
| Minimum detectable effect (experiments) | Smaller effect to detect → much larger sample |
| Statistical power (experiments) | Higher power (e.g. 80%) → larger sample |
Your baseline conversion rate also feeds in. Because juggling these by hand is error-prone, sample size calculators exist to do the arithmetic: you enter your baseline rate, the minimum detectable effect, your confidence level, and your desired power, and the calculator returns the number of observations, or the traffic per variation, you need. These are standard statistical inputs, not figures specific to any one business. Getting the number wrong, in either direction, has consequences.
Too small vs too large
The danger is asymmetric, but both extremes have costs:
- Too small. Random chance dominates, so the estimate can be far from the truth and swing wildly between samples, and the margin of error is wide even when the headline number looks precise. In an A/B test, low power means it can miss a real improvement (a false negative), and a small sample tempts you to call a test early on a noisy lead that often regresses toward no difference.
- Too large. Far less dangerous, but still costly: more time and money, and in an A/B test, running longer than needed keeps visitors in an inferior variation. There is also a subtle trap, a very large sample can flag a trivially small difference as 'statistically significant', because significance means an effect is real, not that it is big enough to matter.
The fix for both is the same: decide in advance your confidence, your power, and the minimum effect size worth acting on, then collect enough data to detect a meaningful difference and stop. That discipline is nowhere more important than in A/B testing.
Sample size in A/B testing
In A/B testing, sample size is the number of visitors (or sessions) you need in each variation before you can trust the result, and getting it right is central to a valid test. It is determined before the test starts, from four inputs: your current baseline conversion rate, the minimum improvement you want to be able to detect (the minimum detectable effect), your confidence level (commonly 95%), and your statistical power (commonly 80%). A sample size calculator turns those into the number of visitors per variation, and dividing by your daily traffic tells you roughly how long the test must run.
The two rules that follow are simple but frequently broken: reach the pre-calculated sample size before drawing any conclusion, and run for whole weeks so the sample covers your normal mix of weekday and weekend visitors rather than an unusual slice. Calling a test early, before it has gathered enough data, is the classic way to be fooled by noise, because early leads are the least stable. Deciding the sample size up front, and honouring it, is what separates a trustworthy A/B test from a coin flip dressed up as data, and a good platform makes honouring it easy.
Sample size with Omniconvert Explore
Omniconvert Explore is an A/B testing and experimentation platform, and it helps you handle sample size correctly at every stage of a test. Before you launch, thinking in terms of sample size, your baseline rate, the effect you want to detect, and your traffic, tells you whether a test is even feasible in a reasonable time, or whether you should test a higher-traffic page or aim to detect a larger effect.
While the test runs, Explore reports statistical significance, which is the platform's way of telling you whether you have gathered enough evidence to trust the result, so you are not left guessing whether the current lead is real or just noise from too little data. Its segmentation also makes clear when a segment is too small to draw conclusions from, guarding against reading too much into a handful of visitors. Together these features encourage the discipline that sample size demands: decide what you need up front, let the test run until it reaches significance across your normal traffic, and resist calling it early. Across more than 70,000 experiments, with an average uplift of 23.2%, Explore is built to turn an adequately sized test into a decision you can trust.
Stop guessing whether a lead is real. Know when a test has gathered enough data.
See how Omniconvert Explore reports significance →Frequently Asked Questions
Sample size is the number of observations, people, sessions, orders, or other units, included in a sample, usually written as n. When you cannot study a whole population, you study a sample of it, and the sample size is simply how many members of the population are in that sample. It is one of the most important decisions in any study, survey, or experiment, because it governs how much you can trust the result. A larger sample size reduces the effect of random chance, so estimates from a big sample vary less from one sample to the next and sit closer to the true population value; a smaller sample is noisier and can swing widely by luck alone. But bigger is not automatically better in every respect, larger samples cost more time and money to collect, and beyond a certain point the extra precision is not worth the extra cost. The goal is not the largest possible sample but a sample large enough to answer your question reliably. Importantly, sample size controls precision, not representativeness: a big sample drawn badly is still biased, so a good sample must be both unbiased in how it was selected and large enough to be precise.
Sample size matters because it determines how reliable and how precise your results are, and therefore whether you can act on them safely. With a small sample, random chance has a big influence: the result you see could easily be a fluke that would not repeat, so conclusions drawn from it are shaky. As the sample grows, chance variation shrinks, the margin of error around your estimate narrows, and the result becomes something you can trust. In experiments such as A/B tests, sample size is what gives a test the statistical power to detect a real difference when one exists; too small a sample can miss a genuine improvement (a false negative) simply because there was not enough data to see it, while also making any apparent win unreliable. There is a cost to getting it wrong in either direction: too small, and you either miss real effects or chase false ones; needlessly large, and you waste time and money, and in an experiment you keep visitors in a losing variation longer than necessary. Choosing the right sample size up front is what lets a study reach a conclusion that is both trustworthy and efficient.
Several factors together determine the sample size you need, and they trade off against one another. The first is the precision you want: the smaller the margin of error you are willing to accept, the larger the sample you need. The second is the confidence level: being more confident (say 95% rather than 90%) that your interval contains the true value requires a larger sample. The third is the variability of the population: the more spread out or diverse the data, the larger the sample needed to pin down an estimate. In experiments specifically, two more factors dominate: the minimum effect size you want to be able to detect (smaller effects need much larger samples, because a tiny difference is harder to distinguish from noise) and the statistical power you want (the probability of detecting a real effect, commonly set at 80%). Your baseline conversion rate also feeds in. Because juggling these by hand is error-prone, sample size calculators exist to do the arithmetic: you enter your baseline rate, the minimum detectable effect, your confidence level, and your desired power, and the calculator returns the number of observations, or the traffic per variation, you need. These are standard statistical inputs, not figures specific to any one business.
If the sample size is too small, your results are unreliable in ways that are easy to act on by mistake. With too few observations, random chance dominates, so the estimate you get can be far from the truth and can swing dramatically from one sample to the next, the margin of error is wide even if the headline number looks precise. In an A/B test, a sample that is too small has low statistical power, meaning it can easily fail to detect a real improvement that exists (a false negative), so you might discard a genuinely better version simply because there was not enough data to prove it. Worse, small samples make it tempting to 'call' a test early on what looks like a big lead, but early results are the noisiest, and that apparent winner often regresses toward no difference as more data arrives. The practical rule is to decide the required sample size before you start, using the effect you want to detect and your confidence and power, and then wait until you reach it before drawing conclusions. Stopping early because the numbers look good is one of the most common and costly mistakes in experimentation, and an adequate, pre-committed sample size is the defence against it.
Yes, a sample can be larger than you need, and while that is far less dangerous than being too small, it still has costs. The clearest cost is practical: collecting more data than necessary takes more time and money, and in an A/B test it means running the experiment longer than you need to, which keeps some visitors in an inferior variation and delays the moment you can roll the winner out to everyone. There is also a subtler statistical point: with a very large sample, a test can flag a difference as 'statistically significant' even when the difference is so small it has no practical value, because significance measures whether an effect is real, not whether it is big enough to matter. That is why it is good practice to decide in advance not just your confidence and power but the minimum effect size worth acting on, so you collect enough data to detect a meaningful difference and then stop, rather than chasing ever-tinier effects. The aim is efficiency: a sample large enough to answer the question with the precision you need, and no larger.
In A/B testing, sample size is the number of visitors (or sessions) you need in each variation before you can trust the result, and getting it right is central to a valid test. It is determined before the test starts, from four inputs: your current baseline conversion rate, the minimum improvement you want to be able to detect (the minimum detectable effect), your confidence level (commonly 95%), and your statistical power (commonly 80%). A sample size calculator turns those into the number of visitors per variation, and dividing by your daily traffic tells you roughly how long the test must run. The two rules that follow are simple but frequently broken: reach the pre-calculated sample size before drawing any conclusion, and run for whole weeks so the sample covers your normal mix of weekday and weekend visitors rather than an unusual slice. Calling a test early, before it has gathered enough data, is the classic way to be fooled by noise, because early leads are the least stable. Deciding the sample size up front, and honouring it, is what separates a trustworthy A/B test from a coin flip dressed up as data.
Omniconvert Explore is an A/B testing and experimentation platform, and it helps you handle sample size correctly at every stage of a test. Before you launch, thinking in terms of sample size, your baseline rate, the effect you want to detect, and your traffic, tells you whether a test is even feasible in a reasonable time, or whether you should test a higher-traffic page or aim to detect a larger effect. While the test runs, Explore reports statistical significance, which is the platform's way of telling you whether you have gathered enough evidence to trust the result, so you are not left guessing whether the current lead is real or just noise from too little data. Its segmentation also makes clear when a segment is too small to draw conclusions from, guarding against reading too much into a handful of visitors. Together these features encourage the discipline that sample size demands: decide what you need up front, let the test run until it reaches significance across your normal traffic, and resist calling it early. Across more than 70,000 experiments, with an average uplift of 23.2%, Explore is built to turn an adequately sized test into a decision you can trust.
Sample size is one of the least glamorous and most decisive choices in any study: how many observations you gather governs how much you can believe the answer. Too few, and random chance runs the show, wide margins of error, missed real effects, and apparent wins that melt as more data arrives. More, and the picture steadies, until the point where extra data only adds cost. The size you need is not a fixed number but the output of a few clear inputs: the precision and confidence you want, how varied the population is, and, in an experiment, the smallest effect worth detecting and the power to catch it, all standard statistics a calculator can turn into a target. And remember the division of labour: sample size buys precision, not representativeness, so a good sample must be both unbiased in how it was drawn and large enough to be reliable. Decide the number up front, reach it before you conclude, and resist the temptation to call it early, that discipline is what a sound experimentation platform is built to support.
Reach a trustworthy sample size before you call a test with Omniconvert Explore
The most common way to be fooled by a test is to stop it too early. Omniconvert Explore reports statistical significance and segments your results, so you know when a test has gathered enough data to trust, and when a segment is too small to read into.