What Test Duration Is: How Long to Run an A/B Test

First published Jun 11, 2019Updated August 21, 20268 min read
Valentin Radu, Founder and CEO of Omniconvert
Valentin Radu
Founder & CEO, Omniconvert · Author, The CLV Revolution
Published: Jun 11, 2019Updated: Aug 21, 2026
Reviewed by Cristina Stefanova, Head of Content
Quick Answer
Test duration is how long an A/B test runs, from launch until you stop collecting data and decide. The right length isn't an arbitrary calendar choice: it's set by the sample size the test needs (driven by your baseline conversion rate, the smallest effect you want to detect, and your confidence and power settings) divided by your daily eligible traffic, with a floor of at least one to two full business cycles. Running across whole weeks matters because behaviour varies by day, weekends often differ from weekdays, so a test that ends mid-week can be skewed. The single most damaging mistake is stopping early: in the first days almost any test shows a false lead that vanishes with more data, and 'peeking' (repeatedly checking and stopping at the first significant-looking moment) inflates the false-positive rate well above the nominal 5%. The discipline: decide the required sample and minimum duration before launch, then run to it and judge by statistical significance. Omniconvert Explore reports significance and confidence as a test runs, so you know when a result is genuinely conclusive, across 70,000+ experiments.
Key Takeaways
  • Test duration is how long an A/B test runs before you decide; the right length is set by the required sample size ÷ daily traffic, with a floor of at least one to two full business cycles.
  • Sample size is driven by your baseline conversion rate, the smallest effect you want to detect, and your confidence (usually 95%) and power (usually 80%) settings; lower base rates and smaller effects need more time.
  • Run across WHOLE weeks (ending on the same weekday you started) so weekend and weekday behaviour are equally represented, a test ended mid-cycle reads an unrepresentative slice of days.
  • The most damaging mistake is stopping early: in the first days almost any test shows a false lead that vanishes with more data, and 'peeking' inflates the false-positive rate above 5%.
  • The discipline: decide the required sample and minimum duration BEFORE launch, then run to it and judge by statistical significance, don't react to interim swings.
7,000+ websites 15+ industries 70,000+ experiments 23.2% avg uplift

The most expensive mistake in A/B testing isn't a bad idea, it's a good test called too soon. In the first few days almost any experiment will show one version leaping ahead, and it's tempting to declare victory and move on. Do that and you're usually shipping random noise. Test duration is the discipline that prevents it: the length of time a test must run to produce a result you can actually trust. This guide explains what test duration is, what determines it, why a test needs to run across full business cycles, why stopping early is so dangerous, and how to set the right duration, drawing on the experimentation practice behind 70,000+ experiments across 7,000+ websites in 15+ industries [CROBenchmark Report 2026, Omniconvert].

One idea runs through it: duration is a calculation with a floor, sample size divided by traffic, never shorter than one full business cycle.

What test duration is

Test duration is how long an A/B test runs, from launch until you stop collecting data and decide. It's not an arbitrary calendar choice: the right duration is the time needed to gather enough data, a large enough sample, across a representative span of days, for a trustworthy result. Two things set it. First, the sample size needed to detect the effect you care about with statistical confidence (driven by your baseline conversion rate, the difference you want to detect, and your confidence and power levels). Second, your traffic: required sample ÷ daily eligible visitors ≈ the days needed. On top, a test should cover full business cycles, typically at least one to two weeks, so normal variation by day of week and marketing activity is included. In short: test duration is set by sample size and traffic, with a floor of at least one full business cycle.

Test duration is the length of time an A/B test runs, from launch until you stop collecting data and make a decision. It is not an arbitrary calendar choice: the right duration is the time needed for the test to gather enough data, a large enough sample, across a representative span of days, to produce a trustworthy result.

Two things determine it: the sample size the test needs to detect the effect you care about, and your traffic, which converts that sample into a number of days. On top of both sits a floor: a test should run long enough to cover full business cycles. That gives a practical answer to the question everyone asks.

How long to run an A/B test

Long enough to reach the sample size your test needs, and never shorter than one full business cycle, for most sites at least one to two weeks. No single magic number fits every test, because the right length depends on traffic and the effect size you're detecting. Work out the sample size to detect the smallest difference you care about at your chosen confidence (commonly 95%) and power (commonly 80%), then divide by daily eligible traffic to estimate the days. Apply two floors: run at least one to two full weeks to capture day-of-week variation, and run across whole weeks (not partway through one) so every day is equally represented. Note a ceiling too: running much longer than needed isn't free, very long tests are more exposed to external changes and to cookie deletion that muddies who's in which group. The goal is enough time to be confident, and no more.

The honest answer is: long enough to reach the sample size your test needs, and never shorter than one full business cycle, which for most sites means at least one to two weeks. There is no single magic number of days that fits every test, because the right length depends on your traffic and the size of the effect you're trying to detect.

Work out, in advance, the sample size required to detect the smallest difference you care about at your chosen confidence (commonly 95%) and power (commonly 80%), then divide that by your daily eligible traffic to estimate the days needed. Whatever that calculation gives you, apply two floors: run for at least one to two full weeks, and run across whole weeks rather than stopping partway through one. There's also a ceiling, running much longer than needed exposes the test to external changes and cookie deletion. But the single biggest error is at the other end: stopping too soon.

Why stopping too early causes problems

Stopping too early is the most common way to reach a wrong conclusion, because early results are dominated by random noise. In the first hours and days, each version has a small sample, so its conversion rate swings widely by chance, one version almost always looks ahead at some point, even if the two are truly identical. Stop the moment a variation looks like a winner and you're likely acting on noise, and the "win" often evaporates or reverses once shipped. Peeking makes it worse: repeatedly checking and stopping at the first moment it crosses significance gives randomness another chance to produce a false positive each time, so continuous peeking inflates the real false-positive rate well above the nominal 5%. The fix is simple but strict: decide the required sample size and minimum duration before launch, and run to that point regardless of how tempting the interim numbers look.

Stopping a test too early is the most common way to reach a wrong conclusion, and it happens because early results are dominated by random noise. In the first hours and days of a test, each version has only a small sample, so its measured conversion rate can swing widely purely by chance, one version will almost always look ahead at some point, even if the two are truly identical.

This is made worse by "peeking": repeatedly checking a running test and stopping at the first moment it crosses the significance threshold. Every time you peek and could stop, you give randomness another chance to produce a false positive, so continuous peeking dramatically inflates the real false-positive rate above the nominal 5%. The table sets out what each duration outcome does to the result.

Source: Omniconvert. How test duration affects the trustworthiness of the result.
Duration outcome What it does to the result Main risk
Too short (below the sample) Reads a small, noisy sample where one version almost always looks ahead by chance False positive, the "win" evaporates or reverses once shipped
Ended mid-cycle Captures an unrepresentative slice of days (e.g. no weekend) Result skewed by day-of-week variation, may not hold in the real week
Right length (sample reached, whole weeks) Enough data across a full business cycle to be statistically confident Trustworthy, this is the target
Much longer than needed Keeps running well past the required sample More exposure to external changes and cookie deletion that muddies the groups

Avoiding the two failure modes at the top means committing to a duration in advance, and that duration has to respect a natural rhythm in your data.

The business cycle

A business cycle here is the natural repeating pattern of customer behaviour over time, most importantly the weekly rhythm, but also patterns tied to paydays, promotions, or seasons. It matters because behaviour is rarely uniform across days: many sites see different conversion rates on weekends vs weekdays, so a sample from only part of the cycle can be systematically unrepresentative. A test run only Monday-Thursday captures no weekend, and its result may not hold once weekend traffic is included. One full week ensures every day is represented; two weeks is safer, smoothing one-off anomalies and confirming the pattern repeats. Rules: never end a test in under a full week, always run across whole weeks rather than stopping partway through one, and extend across more cycles when your business has strong weekly or monthly swings. Respecting the cycle stops a technically-significant result from being an artefact of WHEN, not what, you tested.

A business cycle, in the context of test duration, is the natural repeating pattern of your customers' behaviour over time, most importantly the weekly rhythm, but also longer patterns tied to paydays, promotions, or seasons. It matters because visitor behaviour is rarely uniform across days: many sites see quite different conversion rates on weekends versus weekdays, so a sample drawn from only part of the cycle can be systematically unrepresentative.

Running for at least one full week ensures every day of the week is represented; running for two weeks is safer still, because it smooths out one-off anomalies and confirms the pattern repeats. The rules that follow are: never end a test in less than a full week, always run across whole weeks, and extend across additional cycles when your business has strong weekly or monthly swings. Put the sample-size maths and the cycle floor together, and setting the duration becomes a simple routine.

How to set the right test duration

A short, disciplined process done before launch: (1) establish your baseline conversion rate for the metric; (2) decide the minimum effect worth acting on; (3) choose confidence (commonly 95%) and power (commonly 80%); (4) run a sample-size (power) calculation to get visitors needed per variation; (5) divide by daily eligible traffic per variation to estimate the days; (6) apply the business-cycle floor, round up to at least one to two full weeks, ending on the same weekday you started. Then commit to that duration and run to it, rather than stopping early because a variation looks ahead. If traffic is low and the test would run very long, that's useful information, test a bigger, bolder change (needs a smaller sample) or focus on higher-traffic pages. The aim: a duration long enough for a trustworthy result, fixed in advance and honoured.

Setting the right test duration is a short, disciplined process you complete before launch:

  1. Establish your baseline. Measure the current conversion rate for the metric you'll test.
  2. Decide the minimum detectable effect. The smallest improvement that would be worth acting on.
  3. Choose confidence and power. Commonly 95% confidence and 80% statistical power.
  4. Calculate the sample size. Feed those inputs into a power calculation to get the visitors each variation needs.
  5. Convert to days. Divide the required sample by your daily eligible traffic per variation.
  6. Apply the cycle floor and commit. Round up to at least one to two full weeks, ending on the same weekday you started, then run to that duration regardless of interim swings.

Then, crucially, honour that duration rather than stopping early because a variation looks ahead. Holding to it is far easier when your platform shows you exactly when a result has become trustworthy.

Getting duration right with Omniconvert Explore

Omniconvert Explore is an A/B testing and experimentation platform designed to keep tests statistically honest, and getting the duration right is central. The core protection is against stopping too early: as a test runs, Explore reports statistical significance and confidence, so you see whether a result is genuinely conclusive or an early, noisy blip that will vanish with more data. That turns "how long should I run this?" into a data-backed decision, run until the result is trustworthy, across a full business cycle, not react to interim swings. Its segmentation adds a consideration: a test with enough data overall usually has far less for any narrow segment, so the segment view helps you judge which segment findings have enough data to trust. And because it runs variations concurrently on live traffic, it captures the same days for every version, so the comparison is fair across whatever cycle the test spans. Across 70,000+ experiments, 23.2% average uplift.

Omniconvert Explore is an A/B testing and experimentation platform designed to keep tests statistically honest, and getting the duration right is central to that. The core protection it provides is against the biggest duration mistake, stopping too early: as a test runs, Explore reports statistical significance and confidence, so you can see whether a result is genuinely conclusive or just an early, noisy blip that will vanish with more data.

Its segmentation adds a further duration consideration: a test that has enough data for a reliable overall result usually has far less for any single narrow segment, so Explore's segment view helps you judge which segment findings have accumulated enough data to trust. And because it runs variations concurrently on live traffic, it captures the same days for every version, so the comparison is fair across whatever business cycle the test spans. Across more than 70,000 experiments, with an average uplift of 23.2%, this discipline, run long enough, judge by significance, and don't stop early, is what makes Explore's results dependable.

Know exactly when a result has become trustworthy, and not a moment sooner.

See how Omniconvert Explore reports significance as tests run →

Frequently Asked Questions

1What is test duration?

Test duration is the length of time an A/B test (or any online experiment) runs, from launch until you stop collecting data and make a decision. It is not an arbitrary calendar choice: the right duration is the time needed for the test to gather enough data, a large enough sample, across a representative span of days, to produce a trustworthy result. Two things determine it. The first is the sample size the test needs to detect the effect you care about with statistical confidence, which depends on your baseline conversion rate, the size of the difference you want to detect, and your chosen confidence and power levels. The second is your traffic: dividing the required sample size by how many eligible visitors the tested page receives per day tells you roughly how many days it will take. On top of that, a test should run long enough to cover full business cycles, typically at least one to two weeks, so normal variation by day of week and marketing activity is included. In short: test duration is set by sample size and traffic, with a floor of at least one full business cycle.

2How long should you run an A/B test?

Long enough to reach the sample size your test needs, and never shorter than one full business cycle, which for most sites means at least one to two weeks. There's no single magic number of days, because the right length depends on your traffic and the size of the effect you're detecting. Work out, in advance, the sample size required to detect the smallest difference you care about at your chosen confidence (commonly 95%) and power (commonly 80%), then divide that by your daily eligible traffic to estimate the days. Whatever that gives you, apply two floors. First, run for at least one to two full weeks so the test captures normal day-of-week variation, weekend behaviour often differs from weekday. Second, run across whole weeks rather than stopping partway through one, so every day is equally represented. And note a ceiling: running much longer than needed isn't free, because very long tests are more exposed to external changes and to cookie deletion that muddies who's in which group. The goal is enough time to be confident, and no more.

3Why does stopping a test too early cause problems?

Stopping too early is the most common way to reach a wrong conclusion, because early results are dominated by random noise. In the first hours and days, each version has only a small sample, so its measured conversion rate can swing widely by chance, one version will almost always look ahead at some point, even if the two are truly identical. If you stop the moment a variation looks like a winner, you're very likely acting on that noise rather than a real effect, and the 'win' often evaporates or reverses when you ship it. This is made worse by 'peeking': repeatedly checking a running test and stopping at the first moment it crosses the significance threshold. Every time you peek and could stop, you give randomness another chance to produce a false positive, so continuous peeking dramatically inflates the real false-positive rate above the nominal 5%. The discipline that prevents this is simple but strict: decide the required sample size and minimum duration before you launch, and run to that point regardless of how tempting the interim numbers look.

4What determines how long a test needs to run?

Four factors, and they interact. First, your baseline conversion rate: lower base rates need larger samples (and more time) because conversions are rarer and noisier. Second, the minimum effect you want to detect: the smaller the difference you care about catching, the larger the sample and longer the test, detecting a subtle 2% relative lift takes far more data than an obvious 20% one. Third, your confidence and power settings: demanding higher confidence or higher power both increase the required sample and duration. Together those three determine the sample size. The fourth factor converts that sample into days: your traffic volume, how many eligible visitors reach the tested page per day and how the split is allocated. A high-traffic page reaches the required sample in days; a low-traffic page may take weeks. Overlaying all of it is the business-cycle floor: no matter what the sample-size maths says, run for at least one to two full weeks so the result isn't skewed by an unrepresentative slice of days.

5What is a business cycle in the context of test duration?

A business cycle, here, is the natural repeating pattern of your customers' behaviour over time, most importantly the weekly rhythm, but also longer patterns tied to paydays, promotions, or seasons. It matters because visitor behaviour is rarely uniform across days: many sites see quite different conversion rates on weekends versus weekdays, so a sample drawn from only part of the cycle can be systematically unrepresentative. If a test runs only Monday to Thursday, it captures none of the weekend, and its result may not hold once weekend traffic is included. Running for at least one full week ensures every day is represented; two weeks is safer still, because it smooths out one-off anomalies (a single unusual day, a one-day promotion) and confirms the pattern repeats. The rules that follow: never end a test in less than a full week, always run across whole weeks rather than stopping partway through one, and extend across additional cycles when your business has strong weekly or monthly swings. Respecting the business cycle is what stops a technically-significant result from being an artefact of when, rather than what, you tested.

6How do you set the right test duration?

It's a short, disciplined process you complete before launch. First, establish your baseline conversion rate for the metric you'll test. Second, decide the minimum effect you want to detect, the smallest improvement worth acting on. Third, choose your confidence level (commonly 95%) and statistical power (commonly 80%). Fourth, use those in a sample-size (power) calculation to get the visitors each variation needs. Fifth, divide that required sample by your daily eligible traffic per variation to estimate the days. Sixth, apply the business-cycle floor: round up so the test runs across at least one to two full weeks, ending on the same weekday it started. Then commit to that duration and run to it, rather than stopping early because a variation looks ahead. If traffic is low and the test would need to run very long, that's useful information, it may mean testing a bigger, bolder change (which needs a smaller sample to detect) or focusing on higher-traffic pages. The aim is a duration long enough for a trustworthy result, fixed in advance and honoured.

7How does Omniconvert Explore help you get test duration right?

Omniconvert Explore is an A/B testing and experimentation platform designed to keep tests statistically honest, and getting the duration right is central to that. The core protection is against the biggest duration mistake, stopping too early: as a test runs, Explore reports statistical significance and confidence, so you can see whether a result is genuinely conclusive or just an early, noisy blip that will vanish with more data. That turns 'how long should I run this?' from a guess into a data-backed decision, you run until the result is trustworthy, across a full business cycle, rather than reacting to interim swings. Its segmentation adds a consideration: a test with enough data for a reliable overall result usually has far less for any single narrow segment, so the segment view helps you judge which segment findings have accumulated enough data to trust. And because it runs variations concurrently on live traffic, it captures the same days for every version, so the comparison is fair across whatever business cycle the test spans. Across more than 70,000 experiments, with an average uplift of 23.2%, this discipline is what makes Explore's results dependable.

The takeaway

Test duration is not a calendar guess, it's a calculation with a floor. The calculation comes from the sample size your test needs (driven by your baseline conversion rate, the smallest effect you want to detect, and your confidence and power settings) divided by your daily traffic. The floor is the business cycle: run for at least one to two full weeks, across whole weeks, so weekend and weekday behaviour are both represented and no single odd day skews the result. Get either wrong and the test lies to you, too short and you're reading random noise, ended mid-cycle and you're reading an unrepresentative slice of days. The single most damaging mistake is stopping early because a variation looks ahead, because in the first days almost any test will show a false lead that vanishes with more data. The discipline that prevents all of this is the same one that underpins trustworthy experimentation: decide the required sample size and minimum duration before launch, then run to it and judge by statistical significance, which is exactly what Omniconvert Explore is built to make routine.

Valentin Radu, Founder and CEO of Omniconvert
Founder & CEO, Omniconvert
Valentin Radu is the founder and CEO of Omniconvert. He is an entrepreneur, data-driven marketer, CRO expert, CVO evangelist, international speaker, father, husband, and pet guardian. Valentin is also an Instructor at the Customer Value Optimization (CVO) Academy, an educational project that aims to help companies understand and improve Customer Lifetime Value.

The biggest duration mistake is stopping early on a noisy lead. See how Omniconvert Explore reports significance and confidence as a test runs, so you know when a result is genuinely conclusive rather than a blip.

See Omniconvert Explore →

Run every test long enough to trust it, with Omniconvert Explore

The biggest duration mistake is stopping early on a noisy lead. Omniconvert Explore reports significance and confidence as a test runs, so you know when a result is genuinely conclusive rather than a blip, and segment to see which findings have the data behind them.