What Test Duration Is: How Long to Run an A/B Test
- Test duration is how long an A/B test runs before you decide; the right length is set by the required sample size ÷ daily traffic, with a floor of at least one to two full business cycles.
- Sample size is driven by your baseline conversion rate, the smallest effect you want to detect, and your confidence (usually 95%) and power (usually 80%) settings; lower base rates and smaller effects need more time.
- Run across WHOLE weeks (ending on the same weekday you started) so weekend and weekday behaviour are equally represented, a test ended mid-cycle reads an unrepresentative slice of days.
- The most damaging mistake is stopping early: in the first days almost any test shows a false lead that vanishes with more data, and 'peeking' inflates the false-positive rate above 5%.
- The discipline: decide the required sample and minimum duration BEFORE launch, then run to it and judge by statistical significance, don't react to interim swings.
The most expensive mistake in A/B testing isn't a bad idea, it's a good test called too soon. In the first few days almost any experiment will show one version leaping ahead, and it's tempting to declare victory and move on. Do that and you're usually shipping random noise. Test duration is the discipline that prevents it: the length of time a test must run to produce a result you can actually trust. This guide explains what test duration is, what determines it, why a test needs to run across full business cycles, why stopping early is so dangerous, and how to set the right duration, drawing on the experimentation practice behind 70,000+ experiments across 7,000+ websites in 15+ industries [CROBenchmark Report 2026, Omniconvert].
One idea runs through it: duration is a calculation with a floor, sample size divided by traffic, never shorter than one full business cycle.
What test duration is
Test duration is the length of time an A/B test runs, from launch until you stop collecting data and make a decision. It is not an arbitrary calendar choice: the right duration is the time needed for the test to gather enough data, a large enough sample, across a representative span of days, to produce a trustworthy result.
Two things determine it: the sample size the test needs to detect the effect you care about, and your traffic, which converts that sample into a number of days. On top of both sits a floor: a test should run long enough to cover full business cycles. That gives a practical answer to the question everyone asks.
How long to run an A/B test
The honest answer is: long enough to reach the sample size your test needs, and never shorter than one full business cycle, which for most sites means at least one to two weeks. There is no single magic number of days that fits every test, because the right length depends on your traffic and the size of the effect you're trying to detect.
Work out, in advance, the sample size required to detect the smallest difference you care about at your chosen confidence (commonly 95%) and power (commonly 80%), then divide that by your daily eligible traffic to estimate the days needed. Whatever that calculation gives you, apply two floors: run for at least one to two full weeks, and run across whole weeks rather than stopping partway through one. There's also a ceiling, running much longer than needed exposes the test to external changes and cookie deletion. But the single biggest error is at the other end: stopping too soon.
Why stopping too early causes problems
Stopping a test too early is the most common way to reach a wrong conclusion, and it happens because early results are dominated by random noise. In the first hours and days of a test, each version has only a small sample, so its measured conversion rate can swing widely purely by chance, one version will almost always look ahead at some point, even if the two are truly identical.
This is made worse by "peeking": repeatedly checking a running test and stopping at the first moment it crosses the significance threshold. Every time you peek and could stop, you give randomness another chance to produce a false positive, so continuous peeking dramatically inflates the real false-positive rate above the nominal 5%. The table sets out what each duration outcome does to the result.
| Duration outcome | What it does to the result | Main risk |
|---|---|---|
| Too short (below the sample) | Reads a small, noisy sample where one version almost always looks ahead by chance | False positive, the "win" evaporates or reverses once shipped |
| Ended mid-cycle | Captures an unrepresentative slice of days (e.g. no weekend) | Result skewed by day-of-week variation, may not hold in the real week |
| Right length (sample reached, whole weeks) | Enough data across a full business cycle to be statistically confident | Trustworthy, this is the target |
| Much longer than needed | Keeps running well past the required sample | More exposure to external changes and cookie deletion that muddies the groups |
Avoiding the two failure modes at the top means committing to a duration in advance, and that duration has to respect a natural rhythm in your data.
The business cycle
A business cycle, in the context of test duration, is the natural repeating pattern of your customers' behaviour over time, most importantly the weekly rhythm, but also longer patterns tied to paydays, promotions, or seasons. It matters because visitor behaviour is rarely uniform across days: many sites see quite different conversion rates on weekends versus weekdays, so a sample drawn from only part of the cycle can be systematically unrepresentative.
Running for at least one full week ensures every day of the week is represented; running for two weeks is safer still, because it smooths out one-off anomalies and confirms the pattern repeats. The rules that follow are: never end a test in less than a full week, always run across whole weeks, and extend across additional cycles when your business has strong weekly or monthly swings. Put the sample-size maths and the cycle floor together, and setting the duration becomes a simple routine.
How to set the right test duration
Setting the right test duration is a short, disciplined process you complete before launch:
- Establish your baseline. Measure the current conversion rate for the metric you'll test.
- Decide the minimum detectable effect. The smallest improvement that would be worth acting on.
- Choose confidence and power. Commonly 95% confidence and 80% statistical power.
- Calculate the sample size. Feed those inputs into a power calculation to get the visitors each variation needs.
- Convert to days. Divide the required sample by your daily eligible traffic per variation.
- Apply the cycle floor and commit. Round up to at least one to two full weeks, ending on the same weekday you started, then run to that duration regardless of interim swings.
Then, crucially, honour that duration rather than stopping early because a variation looks ahead. Holding to it is far easier when your platform shows you exactly when a result has become trustworthy.
Getting duration right with Omniconvert Explore
Omniconvert Explore is an A/B testing and experimentation platform designed to keep tests statistically honest, and getting the duration right is central to that. The core protection it provides is against the biggest duration mistake, stopping too early: as a test runs, Explore reports statistical significance and confidence, so you can see whether a result is genuinely conclusive or just an early, noisy blip that will vanish with more data.
Its segmentation adds a further duration consideration: a test that has enough data for a reliable overall result usually has far less for any single narrow segment, so Explore's segment view helps you judge which segment findings have accumulated enough data to trust. And because it runs variations concurrently on live traffic, it captures the same days for every version, so the comparison is fair across whatever business cycle the test spans. Across more than 70,000 experiments, with an average uplift of 23.2%, this discipline, run long enough, judge by significance, and don't stop early, is what makes Explore's results dependable.
Know exactly when a result has become trustworthy, and not a moment sooner.
See how Omniconvert Explore reports significance as tests run →Frequently Asked Questions
Test duration is the length of time an A/B test (or any online experiment) runs, from launch until you stop collecting data and make a decision. It is not an arbitrary calendar choice: the right duration is the time needed for the test to gather enough data, a large enough sample, across a representative span of days, to produce a trustworthy result. Two things determine it. The first is the sample size the test needs to detect the effect you care about with statistical confidence, which depends on your baseline conversion rate, the size of the difference you want to detect, and your chosen confidence and power levels. The second is your traffic: dividing the required sample size by how many eligible visitors the tested page receives per day tells you roughly how many days it will take. On top of that, a test should run long enough to cover full business cycles, typically at least one to two weeks, so normal variation by day of week and marketing activity is included. In short: test duration is set by sample size and traffic, with a floor of at least one full business cycle.
Long enough to reach the sample size your test needs, and never shorter than one full business cycle, which for most sites means at least one to two weeks. There's no single magic number of days, because the right length depends on your traffic and the size of the effect you're detecting. Work out, in advance, the sample size required to detect the smallest difference you care about at your chosen confidence (commonly 95%) and power (commonly 80%), then divide that by your daily eligible traffic to estimate the days. Whatever that gives you, apply two floors. First, run for at least one to two full weeks so the test captures normal day-of-week variation, weekend behaviour often differs from weekday. Second, run across whole weeks rather than stopping partway through one, so every day is equally represented. And note a ceiling: running much longer than needed isn't free, because very long tests are more exposed to external changes and to cookie deletion that muddies who's in which group. The goal is enough time to be confident, and no more.
Stopping too early is the most common way to reach a wrong conclusion, because early results are dominated by random noise. In the first hours and days, each version has only a small sample, so its measured conversion rate can swing widely by chance, one version will almost always look ahead at some point, even if the two are truly identical. If you stop the moment a variation looks like a winner, you're very likely acting on that noise rather than a real effect, and the 'win' often evaporates or reverses when you ship it. This is made worse by 'peeking': repeatedly checking a running test and stopping at the first moment it crosses the significance threshold. Every time you peek and could stop, you give randomness another chance to produce a false positive, so continuous peeking dramatically inflates the real false-positive rate above the nominal 5%. The discipline that prevents this is simple but strict: decide the required sample size and minimum duration before you launch, and run to that point regardless of how tempting the interim numbers look.
Four factors, and they interact. First, your baseline conversion rate: lower base rates need larger samples (and more time) because conversions are rarer and noisier. Second, the minimum effect you want to detect: the smaller the difference you care about catching, the larger the sample and longer the test, detecting a subtle 2% relative lift takes far more data than an obvious 20% one. Third, your confidence and power settings: demanding higher confidence or higher power both increase the required sample and duration. Together those three determine the sample size. The fourth factor converts that sample into days: your traffic volume, how many eligible visitors reach the tested page per day and how the split is allocated. A high-traffic page reaches the required sample in days; a low-traffic page may take weeks. Overlaying all of it is the business-cycle floor: no matter what the sample-size maths says, run for at least one to two full weeks so the result isn't skewed by an unrepresentative slice of days.
A business cycle, here, is the natural repeating pattern of your customers' behaviour over time, most importantly the weekly rhythm, but also longer patterns tied to paydays, promotions, or seasons. It matters because visitor behaviour is rarely uniform across days: many sites see quite different conversion rates on weekends versus weekdays, so a sample drawn from only part of the cycle can be systematically unrepresentative. If a test runs only Monday to Thursday, it captures none of the weekend, and its result may not hold once weekend traffic is included. Running for at least one full week ensures every day is represented; two weeks is safer still, because it smooths out one-off anomalies (a single unusual day, a one-day promotion) and confirms the pattern repeats. The rules that follow: never end a test in less than a full week, always run across whole weeks rather than stopping partway through one, and extend across additional cycles when your business has strong weekly or monthly swings. Respecting the business cycle is what stops a technically-significant result from being an artefact of when, rather than what, you tested.
It's a short, disciplined process you complete before launch. First, establish your baseline conversion rate for the metric you'll test. Second, decide the minimum effect you want to detect, the smallest improvement worth acting on. Third, choose your confidence level (commonly 95%) and statistical power (commonly 80%). Fourth, use those in a sample-size (power) calculation to get the visitors each variation needs. Fifth, divide that required sample by your daily eligible traffic per variation to estimate the days. Sixth, apply the business-cycle floor: round up so the test runs across at least one to two full weeks, ending on the same weekday it started. Then commit to that duration and run to it, rather than stopping early because a variation looks ahead. If traffic is low and the test would need to run very long, that's useful information, it may mean testing a bigger, bolder change (which needs a smaller sample to detect) or focusing on higher-traffic pages. The aim is a duration long enough for a trustworthy result, fixed in advance and honoured.
Omniconvert Explore is an A/B testing and experimentation platform designed to keep tests statistically honest, and getting the duration right is central to that. The core protection is against the biggest duration mistake, stopping too early: as a test runs, Explore reports statistical significance and confidence, so you can see whether a result is genuinely conclusive or just an early, noisy blip that will vanish with more data. That turns 'how long should I run this?' from a guess into a data-backed decision, you run until the result is trustworthy, across a full business cycle, rather than reacting to interim swings. Its segmentation adds a consideration: a test with enough data for a reliable overall result usually has far less for any single narrow segment, so the segment view helps you judge which segment findings have accumulated enough data to trust. And because it runs variations concurrently on live traffic, it captures the same days for every version, so the comparison is fair across whatever business cycle the test spans. Across more than 70,000 experiments, with an average uplift of 23.2%, this discipline is what makes Explore's results dependable.
Test duration is not a calendar guess, it's a calculation with a floor. The calculation comes from the sample size your test needs (driven by your baseline conversion rate, the smallest effect you want to detect, and your confidence and power settings) divided by your daily traffic. The floor is the business cycle: run for at least one to two full weeks, across whole weeks, so weekend and weekday behaviour are both represented and no single odd day skews the result. Get either wrong and the test lies to you, too short and you're reading random noise, ended mid-cycle and you're reading an unrepresentative slice of days. The single most damaging mistake is stopping early because a variation looks ahead, because in the first days almost any test will show a false lead that vanishes with more data. The discipline that prevents all of this is the same one that underpins trustworthy experimentation: decide the required sample size and minimum duration before launch, then run to it and judge by statistical significance, which is exactly what Omniconvert Explore is built to make routine.
Run every test long enough to trust it, with Omniconvert Explore
The biggest duration mistake is stopping early on a noisy lead. Omniconvert Explore reports significance and confidence as a test runs, so you know when a result is genuinely conclusive rather than a blip, and segment to see which findings have the data behind them.