What Sample Size Is: Definition, Why It Matters & How to Set It

First published Feb 14, 2025Updated August 21, 202610 min read
Valentin Radu, Founder and CEO of Omniconvert
Valentin Radu
Founder & CEO, Omniconvert · Author, The CLV Revolution
Published: Feb 14, 2025Updated: Aug 21, 2026
Reviewed by Cristina Stefanova, Head of Content
Quick Answer
Sample size is the number of observations (people, sessions, orders) included in a sample, usually written as n. When you cannot study a whole population, you study a sample, and the sample size is how many members it contains. It governs how much you can trust the result: a larger sample reduces the effect of random chance, so estimates vary less and sit closer to the true value, while a small sample is noisy and can swing by luck alone. But bigger is not automatically better, beyond a point the extra precision is not worth the cost. The size you need depends on the precision (margin of error) and confidence you want, the variability of the population, and, in experiments, the minimum effect you want to detect and the statistical power, all standard inputs a sample size calculator turns into a target. Crucially, sample size buys precision, not representativeness: a big sample drawn badly is still biased. In A/B testing, decide the size up front and reach it before concluding, don't call a test early. Omniconvert Explore reports significance across 70,000+ experiments, 23.2% average uplift.
Key Takeaways
  • Sample size is the number of observations (n) in a sample; it governs how much random chance affects your result and therefore how much you can trust it.
  • A larger sample reduces chance variation and narrows the margin of error, but beyond a point the extra precision isn't worth the extra cost.
  • The size you need depends on desired precision and confidence, population variability, and, in experiments, the minimum detectable effect and statistical power, standard inputs a calculator turns into a target.
  • Sample size controls precision, not representativeness: a large sample drawn badly is still biased, so a good sample must be both unbiased and large enough.
  • In A/B testing, decide the sample size up front and reach it before concluding, calling a test early on a noisy lead is a classic, costly mistake.
7,000+ websites 15+ industries 70,000+ experiments 23.2% avg uplift

Almost every number you report comes from a sample, and how many observations that sample holds quietly decides how much the number can be believed. Sample size is that count. Too few, and random chance runs the show; enough, and the picture steadies into something you can act on. It is one of the least glamorous decisions in any study and one of the most consequential. This guide explains what sample size is, why it matters, what determines the size you need, the problems at both extremes, and its role in A/B testing, and how Omniconvert Explore supports the discipline it demands, drawing on 70,000+ experiments across 7,000+ websites in 15+ industries [CROBenchmark Report 2026, Omniconvert].

One idea holds it together: sample size buys precision, not representativeness, and the right size is the smallest one that answers your question reliably.

What sample size is

Sample size is the number of observations (people, sessions, orders) included in a sample, usually written as n. When you cannot study a whole population, you study a sample, and the sample size is how many members it contains. It governs how much you can trust the result: a larger sample reduces the effect of random chance, so estimates vary less from sample to sample and sit closer to the true value; a smaller sample is noisier and can swing by luck alone. But bigger is not automatically better, larger samples cost more time and money, and beyond a point the extra precision is not worth it. The goal is not the largest sample but one large enough to answer your question reliably. Crucially, sample size controls precision, not representativeness: a big sample drawn badly is still biased, so a good sample must be both unbiased in selection and large enough to be precise.

Sample size is the number of observations, people, sessions, orders, or other units, included in a sample, usually written as n. When you cannot study a whole population, you study a sample of it, and the sample size is simply how many members of the population are in that sample. It is one of the most important decisions in any study, survey, or experiment, because it governs how much you can trust the result.

A larger sample size reduces the effect of random chance, so estimates from a big sample vary less and sit closer to the true population value; a smaller sample is noisier and can swing widely by luck alone. But bigger is not automatically better, larger samples cost more time and money, and beyond a certain point the extra precision is not worth the extra cost. Importantly, sample size controls precision, not representativeness: a big sample drawn badly is still biased, so a good sample must be both unbiased in how it was selected and large enough to be precise. That precision is exactly why it matters.

Why sample size matters

Sample size matters because it determines how reliable and precise your results are, and whether you can act on them safely. With a small sample, random chance has a big influence, the result could be a fluke that would not repeat, so conclusions are shaky. As the sample grows, chance variation shrinks, the margin of error narrows, and the result becomes trustworthy. In experiments, sample size gives a test the statistical power to detect a real difference; too small a sample can miss a genuine improvement (a false negative) simply for lack of data, while making any apparent win unreliable. Getting it wrong costs both ways: too small, and you miss real effects or chase false ones; needlessly large, and you waste time and money and keep visitors in a losing variation longer. Choosing the right size up front is what makes a conclusion both trustworthy and efficient.

Sample size matters because it determines how reliable and how precise your results are, and therefore whether you can act on them safely. With a small sample, random chance has a big influence: the result you see could easily be a fluke that would not repeat, so conclusions drawn from it are shaky. As the sample grows, chance variation shrinks, the margin of error around your estimate narrows, and the result becomes something you can trust.

In experiments such as A/B tests, sample size is what gives a test the statistical power to detect a real difference when one exists; too small a sample can miss a genuine improvement (a false negative) simply because there was not enough data to see it. There is a cost to getting it wrong in either direction: too small, and you either miss real effects or chase false ones; needlessly large, and you waste time and money. Choosing the right sample size up front is what lets a study reach a conclusion that is both trustworthy and efficient, which raises the question of how you choose it.

What determines the sample size you need

Several factors determine the size you need, and they trade off. Precision: the smaller the margin of error you accept, the larger the sample. Confidence level: more confidence (95% vs 90%) needs a larger sample. Population variability: the more spread out the data, the larger the sample to pin down an estimate. In experiments, two more dominate: the minimum effect size you want to detect (smaller effects need much larger samples, a tiny difference is harder to tell from noise) and the statistical power you want (probability of catching a real effect, commonly 80%). Your baseline conversion rate feeds in too. Because juggling these by hand is error-prone, sample size calculators do the arithmetic: enter baseline rate, minimum detectable effect, confidence, and power, and get the observations, or traffic per variation, you need. These are standard statistical inputs.

Several factors together determine the sample size you need, and they trade off against one another:

Source: Omniconvert. What raises or lowers the sample size a study requires.
Factor Effect on required sample size
Desired precision (margin of error) Smaller margin of error → larger sample
Confidence level Higher confidence (e.g. 95% vs 90%) → larger sample
Population variability More varied / diverse data → larger sample
Minimum detectable effect (experiments) Smaller effect to detect → much larger sample
Statistical power (experiments) Higher power (e.g. 80%) → larger sample

Your baseline conversion rate also feeds in. Because juggling these by hand is error-prone, sample size calculators exist to do the arithmetic: you enter your baseline rate, the minimum detectable effect, your confidence level, and your desired power, and the calculator returns the number of observations, or the traffic per variation, you need. These are standard statistical inputs, not figures specific to any one business. Getting the number wrong, in either direction, has consequences.

Too small vs too large

Too small: random chance dominates, the estimate can be far from the truth and swing wildly, and the margin of error is wide even if the headline number looks precise. In an A/B test, low power means it can miss a real improvement (a false negative), and small samples tempt you to "call" a test early on a noisy lead that often regresses to no difference. Too large: far less dangerous, but still costly, more time and money, a test run longer than needed keeping visitors in an inferior variation, and the subtle trap that a very large sample can flag a trivially small difference as "significant" (significance means real, not big enough to matter). The fix for both is to decide in advance your confidence, power, and the minimum effect worth acting on, then collect enough to detect a meaningful difference and stop.

The danger is asymmetric, but both extremes have costs:

  • Too small. Random chance dominates, so the estimate can be far from the truth and swing wildly between samples, and the margin of error is wide even when the headline number looks precise. In an A/B test, low power means it can miss a real improvement (a false negative), and a small sample tempts you to call a test early on a noisy lead that often regresses toward no difference.
  • Too large. Far less dangerous, but still costly: more time and money, and in an A/B test, running longer than needed keeps visitors in an inferior variation. There is also a subtle trap, a very large sample can flag a trivially small difference as 'statistically significant', because significance means an effect is real, not that it is big enough to matter.

The fix for both is the same: decide in advance your confidence, your power, and the minimum effect size worth acting on, then collect enough data to detect a meaningful difference and stop. That discipline is nowhere more important than in A/B testing.

Sample size in A/B testing

In A/B testing, sample size is the number of visitors (or sessions) you need in each variation before you can trust the result. It is set before the test starts, from four inputs: your baseline conversion rate, the minimum improvement to detect (minimum detectable effect), your confidence level (commonly 95%), and your power (commonly 80%). A calculator turns those into visitors per variation, and dividing by daily traffic tells you roughly how long to run. Two rules follow, simple but often broken: reach the pre-calculated size before concluding, and run for whole weeks so the sample covers your normal weekday/weekend mix, not an unusual slice. Calling a test early, before enough data, is the classic way to be fooled by noise, because early leads are least stable. Deciding the size up front and honouring it separates a trustworthy test from a coin flip dressed up as data.

In A/B testing, sample size is the number of visitors (or sessions) you need in each variation before you can trust the result, and getting it right is central to a valid test. It is determined before the test starts, from four inputs: your current baseline conversion rate, the minimum improvement you want to be able to detect (the minimum detectable effect), your confidence level (commonly 95%), and your statistical power (commonly 80%). A sample size calculator turns those into the number of visitors per variation, and dividing by your daily traffic tells you roughly how long the test must run.

The two rules that follow are simple but frequently broken: reach the pre-calculated sample size before drawing any conclusion, and run for whole weeks so the sample covers your normal mix of weekday and weekend visitors rather than an unusual slice. Calling a test early, before it has gathered enough data, is the classic way to be fooled by noise, because early leads are the least stable. Deciding the sample size up front, and honouring it, is what separates a trustworthy A/B test from a coin flip dressed up as data, and a good platform makes honouring it easy.

Sample size with Omniconvert Explore

Omniconvert Explore is an A/B testing and experimentation platform, and it helps you handle sample size at every stage. Before launch, thinking in terms of sample size, baseline rate, effect to detect, and traffic, tells you whether a test is feasible in reasonable time, or whether to test a higher-traffic page or aim for a larger effect. While it runs, Explore reports statistical significance, the platform's way of telling you whether you have gathered enough evidence to trust the result, so you are not guessing whether the lead is real or just noise. Its segmentation also flags when a segment is too small to conclude from, guarding against reading too much into a handful of visitors. Together these encourage the discipline sample size demands: decide up front, run to significance across normal traffic, and resist calling it early. Across 70,000+ experiments, with 23.2% average uplift.

Omniconvert Explore is an A/B testing and experimentation platform, and it helps you handle sample size correctly at every stage of a test. Before you launch, thinking in terms of sample size, your baseline rate, the effect you want to detect, and your traffic, tells you whether a test is even feasible in a reasonable time, or whether you should test a higher-traffic page or aim to detect a larger effect.

While the test runs, Explore reports statistical significance, which is the platform's way of telling you whether you have gathered enough evidence to trust the result, so you are not left guessing whether the current lead is real or just noise from too little data. Its segmentation also makes clear when a segment is too small to draw conclusions from, guarding against reading too much into a handful of visitors. Together these features encourage the discipline that sample size demands: decide what you need up front, let the test run until it reaches significance across your normal traffic, and resist calling it early. Across more than 70,000 experiments, with an average uplift of 23.2%, Explore is built to turn an adequately sized test into a decision you can trust.

Stop guessing whether a lead is real. Know when a test has gathered enough data.

See how Omniconvert Explore reports significance →

Frequently Asked Questions

1What is sample size?

Sample size is the number of observations, people, sessions, orders, or other units, included in a sample, usually written as n. When you cannot study a whole population, you study a sample of it, and the sample size is simply how many members of the population are in that sample. It is one of the most important decisions in any study, survey, or experiment, because it governs how much you can trust the result. A larger sample size reduces the effect of random chance, so estimates from a big sample vary less from one sample to the next and sit closer to the true population value; a smaller sample is noisier and can swing widely by luck alone. But bigger is not automatically better in every respect, larger samples cost more time and money to collect, and beyond a certain point the extra precision is not worth the extra cost. The goal is not the largest possible sample but a sample large enough to answer your question reliably. Importantly, sample size controls precision, not representativeness: a big sample drawn badly is still biased, so a good sample must be both unbiased in how it was selected and large enough to be precise.

2Why does sample size matter?

Sample size matters because it determines how reliable and how precise your results are, and therefore whether you can act on them safely. With a small sample, random chance has a big influence: the result you see could easily be a fluke that would not repeat, so conclusions drawn from it are shaky. As the sample grows, chance variation shrinks, the margin of error around your estimate narrows, and the result becomes something you can trust. In experiments such as A/B tests, sample size is what gives a test the statistical power to detect a real difference when one exists; too small a sample can miss a genuine improvement (a false negative) simply because there was not enough data to see it, while also making any apparent win unreliable. There is a cost to getting it wrong in either direction: too small, and you either miss real effects or chase false ones; needlessly large, and you waste time and money, and in an experiment you keep visitors in a losing variation longer than necessary. Choosing the right sample size up front is what lets a study reach a conclusion that is both trustworthy and efficient.

3What determines the sample size you need?

Several factors together determine the sample size you need, and they trade off against one another. The first is the precision you want: the smaller the margin of error you are willing to accept, the larger the sample you need. The second is the confidence level: being more confident (say 95% rather than 90%) that your interval contains the true value requires a larger sample. The third is the variability of the population: the more spread out or diverse the data, the larger the sample needed to pin down an estimate. In experiments specifically, two more factors dominate: the minimum effect size you want to be able to detect (smaller effects need much larger samples, because a tiny difference is harder to distinguish from noise) and the statistical power you want (the probability of detecting a real effect, commonly set at 80%). Your baseline conversion rate also feeds in. Because juggling these by hand is error-prone, sample size calculators exist to do the arithmetic: you enter your baseline rate, the minimum detectable effect, your confidence level, and your desired power, and the calculator returns the number of observations, or the traffic per variation, you need. These are standard statistical inputs, not figures specific to any one business.

4What happens if the sample size is too small?

If the sample size is too small, your results are unreliable in ways that are easy to act on by mistake. With too few observations, random chance dominates, so the estimate you get can be far from the truth and can swing dramatically from one sample to the next, the margin of error is wide even if the headline number looks precise. In an A/B test, a sample that is too small has low statistical power, meaning it can easily fail to detect a real improvement that exists (a false negative), so you might discard a genuinely better version simply because there was not enough data to prove it. Worse, small samples make it tempting to 'call' a test early on what looks like a big lead, but early results are the noisiest, and that apparent winner often regresses toward no difference as more data arrives. The practical rule is to decide the required sample size before you start, using the effect you want to detect and your confidence and power, and then wait until you reach it before drawing conclusions. Stopping early because the numbers look good is one of the most common and costly mistakes in experimentation, and an adequate, pre-committed sample size is the defence against it.

5Can a sample size be too large?

Yes, a sample can be larger than you need, and while that is far less dangerous than being too small, it still has costs. The clearest cost is practical: collecting more data than necessary takes more time and money, and in an A/B test it means running the experiment longer than you need to, which keeps some visitors in an inferior variation and delays the moment you can roll the winner out to everyone. There is also a subtler statistical point: with a very large sample, a test can flag a difference as 'statistically significant' even when the difference is so small it has no practical value, because significance measures whether an effect is real, not whether it is big enough to matter. That is why it is good practice to decide in advance not just your confidence and power but the minimum effect size worth acting on, so you collect enough data to detect a meaningful difference and then stop, rather than chasing ever-tinier effects. The aim is efficiency: a sample large enough to answer the question with the precision you need, and no larger.

6What is sample size in A/B testing?

In A/B testing, sample size is the number of visitors (or sessions) you need in each variation before you can trust the result, and getting it right is central to a valid test. It is determined before the test starts, from four inputs: your current baseline conversion rate, the minimum improvement you want to be able to detect (the minimum detectable effect), your confidence level (commonly 95%), and your statistical power (commonly 80%). A sample size calculator turns those into the number of visitors per variation, and dividing by your daily traffic tells you roughly how long the test must run. The two rules that follow are simple but frequently broken: reach the pre-calculated sample size before drawing any conclusion, and run for whole weeks so the sample covers your normal mix of weekday and weekend visitors rather than an unusual slice. Calling a test early, before it has gathered enough data, is the classic way to be fooled by noise, because early leads are the least stable. Deciding the sample size up front, and honouring it, is what separates a trustworthy A/B test from a coin flip dressed up as data.

7How does Omniconvert Explore help with sample size?

Omniconvert Explore is an A/B testing and experimentation platform, and it helps you handle sample size correctly at every stage of a test. Before you launch, thinking in terms of sample size, your baseline rate, the effect you want to detect, and your traffic, tells you whether a test is even feasible in a reasonable time, or whether you should test a higher-traffic page or aim to detect a larger effect. While the test runs, Explore reports statistical significance, which is the platform's way of telling you whether you have gathered enough evidence to trust the result, so you are not left guessing whether the current lead is real or just noise from too little data. Its segmentation also makes clear when a segment is too small to draw conclusions from, guarding against reading too much into a handful of visitors. Together these features encourage the discipline that sample size demands: decide what you need up front, let the test run until it reaches significance across your normal traffic, and resist calling it early. Across more than 70,000 experiments, with an average uplift of 23.2%, Explore is built to turn an adequately sized test into a decision you can trust.

The takeaway

Sample size is one of the least glamorous and most decisive choices in any study: how many observations you gather governs how much you can believe the answer. Too few, and random chance runs the show, wide margins of error, missed real effects, and apparent wins that melt as more data arrives. More, and the picture steadies, until the point where extra data only adds cost. The size you need is not a fixed number but the output of a few clear inputs: the precision and confidence you want, how varied the population is, and, in an experiment, the smallest effect worth detecting and the power to catch it, all standard statistics a calculator can turn into a target. And remember the division of labour: sample size buys precision, not representativeness, so a good sample must be both unbiased in how it was drawn and large enough to be reliable. Decide the number up front, reach it before you conclude, and resist the temptation to call it early, that discipline is what a sound experimentation platform is built to support.

Valentin Radu, Founder and CEO of Omniconvert
Founder & CEO, Omniconvert
Valentin Radu is the founder and CEO of Omniconvert. He is an entrepreneur, data-driven marketer, CRO expert, CVO evangelist, international speaker, father, husband, and pet guardian. Valentin is also an Instructor at the Customer Value Optimization (CVO) Academy, an educational project that aims to help companies understand and improve Customer Lifetime Value.

The most common way to be fooled by a test is to stop it too early. See how Omniconvert Explore reports significance and segments results so you reach a trustworthy sample before you call it.

See Omniconvert Explore →

Reach a trustworthy sample size before you call a test with Omniconvert Explore

The most common way to be fooled by a test is to stop it too early. Omniconvert Explore reports statistical significance and segments your results, so you know when a test has gathered enough data to trust, and when a segment is too small to read into.