A/B TestingCRO Strategy

How to Create an A/B Test: A 4-Stage Framework

First published Oct 20, 2019Updated September 16, 202615 min read
Valentin Radu, Founder and CEO of Omniconvert
Valentin Radu
Founder & CEO, Omniconvert · Author, The CLV Revolution
Published: Oct 20, 2019Updated: Sep 16, 2026
Reviewed by Cristina Stefanova, Head of Content
Two gift box versions side by side, with most figurines gathered around the blue box tagged as the winner
Quick Answer
To create an A/B test, follow four stages. First, identify a problem that blocks a business goal, and pick one primary metric to measure it. Second, write a hypothesis that names the change you believe will fix the problem, drawing on data, case studies, best practices, theory and experience. Third, prioritize your ideas, build the variant, calculate the sample size before launch and check that every version renders and tracks correctly. Fourth, let the test reach its planned sample size, read the primary metric, segment the results for new ideas and implement the winner. Omniconvert Explore lets you build variants in a visual editor and run the test without code.
Key Takeaways
  • An A/B testing framework has four stages: identify problems, develop a hypothesis, build and run the test, then analyze results and implement changes.
  • A problem is worth testing when it affects business and website goals; a problem that affects only usability is a UI/UX issue, not a CRO issue.
  • Hypotheses come from four sources: case studies, best practices, theory and experience, and one hypothesis can produce several variants.
  • Sample size should be calculated before launch from the baseline conversion rate, minimum detectable effect, significance level and statistical power.
  • Stopping a fixed-horizon test at the first significant result inflates false positives; stop at the planned sample size or use sequential or Bayesian statistics built for early stopping.
70,000+ experiments run on Explore 23.2% average conversion uplift 7,000+ websites optimized 15+ industries covered

To create an A/B test, follow a four-stage framework. First, identify a problem that affects a business goal. Second, write a hypothesis that explains the problem and proposes a fix. Third, prioritize, build the variant, set the sample size and run the test. Fourth, analyze the result against your primary metric, segment it, record what you learned and implement the winner.

By now you may have done the hard work: dug through your quantitative data and talked to customers to understand their needs. You might even have a few ideas on how to fix what you found. Now comes the fun part, where you test those ideas and see real improvements in your conversion rate. Since most of your tests will be simple split tests (that is, A/B tests), a simple framework helps you run them more efficiently and trust what they tell you.

This guide walks through that framework. If you need the basics first, see what A/B testing is.

Why should you use an A/B testing framework?

An A/B testing framework makes testing systematic instead of random. It connects every test to a business goal, forces a hypothesis before any design work, and fixes the rules for reading results before the data arrives. That is what makes the results trustworthy enough to act on, and what lets each test inform the next one.

In conversion rate optimization (CRO), a structured approach gives you results you can be confident in when you make changes. Two things explain why:

  • CRO is a scientific process. You first identify a problem (such as a poor click-through rate on a button), then propose a solution and test it. A good test does not assume; it checks.
  • CRO is a continuous process. There is no "final" version of a page. Ideally, you keep testing variants to push conversion rates higher.

Because CRO is both scientific and continuous, you will cycle through data collection, brainstorming and testing again and again. A framework keeps those cycles organized. It has four stages:

  1. Identify problems based on business, marketing or conversion goals.
  2. Develop hypotheses from collected qualitative and quantitative data to solve those problems.
  3. Build and test page variants based on the hypothesis.
  4. Analyze the results and implement changes, then feed what you learned into the next test.

Stage 1: How do you identify what to A/B test?

Start from what matters to the business, not from what looks wrong on the page. Define your business objective, the website goal that serves it, the KPI that measures it and the target that defines success. Then look in your data for problems that block those goals. Each test should focus on one primary metric.

The first step is to decide what you want to test and why. Work from broad to specific:

  • Business objectives: the outcome you want, such as more sales, more users or more page views. This depends on the kind of business you run and its stage. A media company might focus on page views, while a new startup might trade profitability for user growth.
  • Website goals: the site behaviors most likely to deliver that objective, such as reducing cart abandonment or adding better product images to drive sales.
  • Key performance indicators (KPIs): measurable metrics that track progress toward the objective. If the goal is more sales, a KPI might be "number of orders processed daily".
  • Target metrics: the level you want the KPI to reach, such as "x sales or sign-ups per month". This is what decides whether your work succeeded.

Your objectives will not always be singular. You might want to grow revenue and newsletter sign-ups at the same time. Each test, however, should have one primary metric.

Once your goals are defined, dig through your data for problems that connect to them. Common warning signs are:

  • High bounce rate on money pages, such as product or payment pages
  • High shopping cart abandonment
  • Low average time on site
  • Low NPS (Net Promoter Score)
  • An issue customers mention often in interviews and surveys, such as a site bug
  • Poor eye-tracking or mouse-tracking results (visitors not looking where you want them to)
  • Poor results from usability tests

For each problem you find, ask:

  • Does this problem affect my long-term business objectives?
  • Does it affect my website goals?
  • Is this a KPI for my business? If yes, what should the target be?
  • Is there a short- or long-term monetary cost attached to it?

If the answer to any of these is "yes", the problem is worth testing and solving.

Example: a SaaS business

Suppose you run a SaaS business that sells marketing software. The objective might be revenue growth; the website goal, more free-trial sign-ups from the pricing page; the KPI, trial sign-ups per week; and the target, a specific weekly number your sales plan depends on. A pricing page with heavy traffic and few sign-ups is then a clear candidate for testing.

CRO problems vs. UI/UX problems

Businesses often confuse CRO issues with UI/UX issues. It is easy to see why: good UI/UX generally leads to better conversions and customer loyalty, so every problem on your site can look like a CRO problem.

In truth, a problem is a CRO problem only if it affects your business and website goals. Otherwise, it is a design problem. Suppose customers email to complain that they cannot find the "Contact Us" button. That will surface in your qualitative research. Unless it affects your business objectives, it remains a usability issue. It is worth fixing, but it should not take focus away from your goals. If, instead, your research shows customers abandon their carts because they do not trust your site, that is a CRO issue: higher abandonment means lower sales.

  • If it affects your business and website goals, it is a CRO issue.
  • If it affects only your site's usability, it is a UI/UX issue.

In your testing, prioritize CRO issues over UI/UX issues.

Stage 2: How do you develop an A/B testing hypothesis?

A hypothesis states the problem, the change you believe will fix it and the metric that will prove it. You build it from four sources: case studies, best practices, theory and experience. One hypothesis can produce several variants, as long as each one is measurable and could plausibly solve the problem.

Suppose your data shows customers abandon their carts at the checkout page. That is a clear CRO issue, so you write a hypothesis:

Hypothesis: Improving trustworthiness, reducing friction and addressing customer FUDs (fears, uncertainties and doubts) will reduce the shopping cart abandonment rate.

From accepted practice, existing case studies and psychology, you can then propose several solutions:

  1. Add more payment methods so customers can pay the way they prefer.
  2. Add trust seals and security badges to signal safety.
  3. Add trust markers such as testimonials and social proof throughout the page.
  4. Offer a 30-day money-back guarantee to build confidence.

Each is a different solution to the same stated problem. For a deeper guide to the format, see what a CRO hypothesis is.

Where hypotheses come from

Developing a hypothesis can be tricky, because you often rely on subjective data and intuition. Four sources help:

Source: Omniconvert
SourceWhat it gives youWhat to watch for
Case studiesSolutions that worked for someone else in a similar situationRelevance, audience overlap, statistical rigor, recency and scope
Best practicesA sensible default that is right in most cases"Most cases" is not "every case"; a best-practice page is a starting point
TheoryPrinciples from psychology, design and copywriting that explain behaviorTheory tells you why something might work, not that it will work on your site
ExperiencePattern recognition from optimizing similar sitesPast wins are ideas to test, not results to assume

1. Case studies

CRO relies heavily on case studies to guide best practices and spark new hypotheses. At the hypothesis stage your job is to collect as many plausible solutions as possible, so case studies are very useful. Start with our case studies and this collection of CRO case studies.

Not every case study applies to you. Check:

  • Relevance: Does it describe a problem like yours? A larger button may have increased sign-ups for a free SaaS app but may not do the same on the checkout page of an eCommerce store.
  • Audience overlap: Does the site have an audience like yours? Results from a sports eCommerce store may not apply to a media company.
  • Accuracy: Was the test run to a planned sample size and a stated significance level? Many published case studies are not statistically sound.
  • Recency: Are the findings recent? Devices, browsing habits and design conventions change, and what worked years ago may not hold today.
  • Scope: Many case studies report "small wins" on micro-conversions, such as a button's click-through rate. Unless the change also moves your target metric, the win is not meaningful for you.

2. Best practices

A best practice is a choice assumed to be correct by default. Placing the navigation bar at the top of the page is one. You could place it in the sidebar or footer to create a unique experience, but for almost every site the top of the page is the default. Best practices come from three places:

  • Convention: some choices are made because that is how things have always been done. Following convention usually helps, because customers already expect it. Breaking it can confuse them.
  • Theory: psychology suggests people equate strong social proof with trustworthiness, so reviews and testimonials became a best practice.
  • Testing: when usability and CRO tests keep producing the same insight, it becomes a best practice. Descriptive CTAs such as "Download eBook" are commonly recommended over generic ones such as "Submit" for this reason.

Best practices hold in most cases, not all. A page that follows them is a good start, not the final version.

3. Theory

CRO sits where sales, design and psychology meet, so your hypotheses will draw on theory from each. Good UI/UX design relies heavily on Gestalt psychology, such as the laws of proximity and similarity. News homepages use these principles to group content: main headlines get large type, section subheadings smaller type, and navigation links a small, uniform size.

CRO also draws on sales and copywriting principles refined over decades, many rooted in psychology, such as scarcity and authority. You do not need to master all of it, but a firm grasp of the psychological and emotional triggers behind purchases will help you write better hypotheses.

4. Experience

Case studies and theory cannot replace experience. A practitioner who has already optimized a site like yours, in the same niche, brings ideas a newcomer would not have. The strongest hypotheses often come from that experience. Treat them the same way as any other idea, though: as something to test.

Bringing it together: one problem, four variants

Problem: a poor click-through rate from the product page to the checkout page. Possible causes include a lack of trust, doubts about product quality, a checkout button that is hard to find, and high or unclear shipping costs.

Hypothesis: Reducing the trust deficit, reassuring customers about product quality and making the page easier to use will improve the product page's conversion rate.

From that one hypothesis, you can brainstorm a variant from each source:

  • Case study: a similar site saw better conversions with a larger, higher-contrast checkout button. Variant A enlarges the button.
  • Best practice: product pages should show reviews and testimonials. Variant B adds social proof.
  • Theory: conversions rise when customers trust the site more. Variant C adds "Featured in" media logos.
  • Experience: on similar sites, a free-shipping threshold has lifted conversions. Variant D adds a free-shipping-over-X message.

Before you move on, check two things: that each variant can be measured against your primary metric, and that each could plausibly solve the problem. Also expect your first test not to hit the target. You will test and refine repeatedly, and you are not limited to one idea per problem.

Stage 3: How do you build and run the A/B test?

Prioritize your ideas, then design the test: one control, one or more variants, one primary metric and a traffic split. Calculate the sample size before launch from your baseline conversion rate and the smallest effect you care about. Build the variants, check they render correctly on your visitors' browsers and devices, confirm tracking works, and launch.

Prioritize what to test first

Businesses often spend heavily testing ideas that barely move the bottom line. Score each idea: low scores for changes you expect to give a minor boost, high scores for those that address a root cause. While scoring, consider:

  • Ease of implementation: the complexity, risk and time involved compared with other ideas.
  • Opportunity: how much the change could move conversions, not just nudge them.
  • Stage 1 goals: whether the variant targets one of the metrics you defined. If not, score it lower.

A simple spreadsheet of ideas and scores is enough. For a fuller method, see how to build an A/B testing plan.

Design the test

The current version of the page is the control. Each changed version is a variant (or challenger). A split test divides traffic between them, usually equally. With a control and three variants on an even split, each version gets one quarter of the traffic.

For example, to test a contact button you might run:

  • Control: the button unchanged
  • Variant A: the button in blue
  • Variant B: the button in red
  • Variant C: the button in orange

Each variant is compared with the control. Every extra variant needs more traffic and raises the chance of a false positive among the comparisons, so add variants only when you have the traffic to support them.

Define the primary metric for the test. It must match what the change is meant to affect: if you want more clicks on a button, track click-through rate, not time on site. Most testing tools also track other metrics, such as engagement. Treat those as secondary metrics that add context; the decision rests on the primary metric.

Set the sample size and duration before you launch

Decide how many visitors each version needs before the test starts. The calculation uses your baseline conversion rate, the minimum detectable effect (the smallest lift worth acting on), the significance level (5% is the common convention, often described as 95% confidence) and the statistical power (80% is common). A lower baseline or a smaller effect means a larger sample. See sample size, statistical power analysis and test duration for the details.

Divide the required sample by your daily traffic to estimate duration, then round up to whole weeks so every day of the week is covered equally. If the estimate runs to several months, test a bolder change, test on a higher-traffic page, or choose a metric closer to the change.

Build the variants

Making page variants once meant hiring a designer and a developer for every version, which made CRO expensive for most businesses. Modern testing tools let you change pages without code. Omniconvert Explore includes a visual editor: make your edits, save the variant and start the test. Changes that affect functionality, such as a new checkout flow, may still need developer help.

Check quality before launch

Before launch, make sure every variant renders correctly on the browsers and devices your visitors actually use. Check the browser and device breakdown in your analytics tool (in Google Analytics 4, the Tech details report), then preview each variant on the main combinations. You can use BrowserStack for this. Your page should render well for the large majority of your traffic, on desktop and mobile.

Then confirm that your tracking code fires and that conversions are recorded for every version. Once both checks pass, launch the test.

Build variants in a visual editor, split traffic and read results with frequentist or Bayesian statistics. Omniconvert Explore has powered 70,000+ experiments across 7,000+ websites.

See Omniconvert Explore →

Stage 4: How do you analyze A/B test results and implement changes?

Let the test reach its planned sample size, then read the primary metric: the control wins, a variant wins, or there is no significant difference. Do not stop early because a result looks significant mid-test, unless your tool uses a sequential method built for that. Segment the results for ideas, record what you learned, implement the winner and plan the next test.

When to stop a test

With a classic fixed-horizon test, you stop when the planned sample size is reached, not when the result first crosses the significance line. Checking results daily and stopping at the first significant reading inflates the false-positive rate well above the 5% you planned for. If you need to monitor results as they come in and stop early, use a tool whose statistics are designed for it, such as a sequential test or a Bayesian approach with a pre-set decision rule. For a full treatment, see statistical significance in A/B testing.

Resources are limited, so most programs do not test one page forever. Testing on a given problem usually ends when you reach your target metric, and your effort moves to the next priority.

Read the result

Whatever you test, the outcome is one of three:

  • Control wins: the original outperforms every variant.
  • Variant wins: one variant beats the control on the primary metric.
  • No result: no statistically significant difference at the planned sample size. This means you could not detect an effect of the size you planned for, not that the versions are proven identical.

Your test also tracks several kinds of metric:

  • Primary metric: the one you are optimizing for, such as new account sign-ups.
  • Secondary metrics: other conversions the tool tracks, such as newsletter sign-ups. Watch them for harm, such as a variant that lifts sign-ups but lowers revenue.
  • Interaction metrics: useful context that is not a conversion, such as time on page or bounce rate.

If you run several tests at the same time, keep them on separate pages or audiences where possible. Overlapping tests on the same page and element can interact, which makes it hard to tell which change caused the result. Tests on unrelated parts of the site can usually run side by side. All versions within one test must run over the same period.

Segment after the test

Segmenting results can reveal a lot about your audience. Say variant A beats the control on overall conversion rate, but among iPhone users the control does better. That is worth knowing, but treat it with care: segments have smaller samples, and the more segments you check, the more likely one shows a difference by chance. Use a segment finding as a hypothesis, and confirm it with a test targeted at that segment before you show different versions to different visitors.

Record what you learned

Keep a log of learnings from every test. It becomes your own database of testing ideas and reveals patterns about your customers. For example, if you test CTA copy and the version that emphasizes "free" wins, you can say your audience responds to free offers, at least on that page. That learning can shape future tests, such as featuring "free" in your landing page copy. Record losing and inconclusive tests too; they tell you what not to repeat.

Test further

Each result opens new tests. If a larger button wins, you can test its color or its position at the new size. This makes the framework open-ended: findings from one test point to the next.

Implement the winner

When a variant wins, make it the new default version of the page. It then becomes the control for your next round of testing, and the cycle starts again at Stage 1.

Frequently Asked Questions

1How do you create an A/B test?

Follow four stages. Identify a problem that affects a business goal. Write a hypothesis that names the change and the metric it should move. Prioritize your ideas, build the variant, calculate the sample size and check that the variant renders and tracks correctly before launch. When the test reaches its planned sample size, analyze the primary metric, implement the winner and record what you learned.

2What should you A/B test first?

Test the problems that block your business goals, not the ones that only look wrong. Score each idea on its potential impact on your target metric, its ease of implementation and whether it addresses a root cause. High-traffic money pages, such as product, pricing and checkout pages, are usually the best starting point because they affect revenue and reach a sample size faster.

3What makes a good A/B testing hypothesis?

A good hypothesis names the problem, the change you believe will fix it and the metric that will show whether it worked. It is based on evidence, such as analytics, surveys, usability tests, case studies or experience, and it is measurable. For example: improving trust signals on the checkout page will reduce cart abandonment.

4How long should an A/B test run?

An A/B test should run until it reaches the sample size you calculated before launch, and for whole weeks so every day of the week is represented. The sample size depends on your baseline conversion rate, the minimum effect you want to detect, the significance level and the statistical power. Low traffic and small effects mean longer tests.

5Can you stop an A/B test as soon as it reaches significance?

Not with a classic fixed-horizon test. Checking repeatedly and stopping at the first significant result inflates the false-positive rate well above the level you planned for. Stop at the planned sample size, or use a tool with sequential or Bayesian statistics designed for continuous monitoring and early stopping.

6What confidence level should an A/B test use?

The common convention is a 5% significance level, often described as 95% confidence, with 80% statistical power. Set both before the test starts. A stricter level lowers the risk of a false positive but needs a larger sample.

7How many variants can an A/B test have?

As many as your traffic supports. Each variant splits the traffic further, so every added variant makes the test longer. It also adds comparisons against the control, which raises the chance that one shows a difference by chance. With limited traffic, test one strong variant against the control.

8What is the difference between a CRO problem and a UI/UX problem?

A problem is a CRO problem if it affects your business and website goals, such as customers abandoning carts because they do not trust the site. A problem that affects only usability, without affecting those goals, is a UI/UX problem. Both are worth fixing, but CRO problems should come first in your testing plan.

Where to start

Start with one problem that blocks a business goal, not with a list of page changes. Write the hypothesis before anyone opens a design tool, and fix the sample size and primary metric before the test goes live. Then let the data finish before you decide. A test that loses or ends without a clear result still teaches you something about your customers, so log it. Run the cycle again with what you learned, and each test will be better aimed than the last.

Valentin Radu, Founder and CEO of Omniconvert
Founder & CEO, Omniconvert
Valentin Radu is the founder and CEO of Omniconvert. He is an entrepreneur, data-driven marketer, CRO expert, CVO evangelist, international speaker, father, husband, and pet guardian. Valentin is also an Instructor at the Customer Value Optimization (CVO) Academy, an educational project that aims to help companies understand and improve Customer Lifetime Value.

Create and run your next A/B test

Omniconvert Explore lets you build variants in a visual editor, split traffic and read results with frequentist or Bayesian statistics. More than 70,000 experiments across 7,000+ websites and 15+ industries, with a 23.2% average conversion uplift.