Split Testing for Pricing: How to Do It Safely
- Revenue per visitor is the primary metric for a pricing test. Conversion rate on its own will recommend the cheaper price almost every time.
- Showing two visitors different prices for the same item at the same moment is the riskiest possible design. Amazon apologized and refunded 6,896 customers after a five-day random DVD price test in September 2000.
- In the EU, the Omnibus Directive (EU) 2019/2161 requires traders to tell consumers when a price has been personalized through automated decision-making. Have counsel review any live price test.
- Safer designs exist and answer most pricing questions: sequential testing, testing new visitors only, and testing price framing, bundles, and free-shipping thresholds rather than the number itself.
- Pricing tests need more traffic than layout tests, because order values vary. Size the test before you launch it, and judge the winner on margin and repeat purchases, not on week-one revenue.
When it comes to pricing, small changes have large effects. Imagine buying concert tickets: the difference between $50 and $60 does not seem like much until you are on the checkout page, deciding whether to hit "buy." That small jump can be the difference between buying now and waiting for a deal.
Finding the right price point takes a data-driven approach, and split testing is how most teams get there. But pricing is not a headline or a button color. A pricing test can charge two people different amounts for the same thing, and that carries consequences a layout test never does. This guide covers what you can test, what you should not, which metric decides the winner, and how to design a test whose result you can actually trust.
What split testing for pricing is
A/B testing compares two versions of something (a page, an ad, a price) to see which performs better. Applied to pricing, it answers a question that no amount of internal debate can settle: what will this audience actually pay?
The mechanics are the same as any other test. Visitors are assigned at random to a control group and one or more variant groups. The control sees the current price. A variant sees a different price, a different bundle, or a different way of presenting the same price. You then compare what each group produced.
What changes with pricing is the stakes. A losing button-color test costs you a week. A losing price test can cost margin on every order placed while it ran, annoy customers who find out, and in some markets attract questions from a regulator. That is why the design deserves more thought than the hypothesis.
The risks nobody warns you about
The cautionary tale is old and still the best one. In September 2000, Amazon ran a five-day test that offered randomly varying discounts on 68 DVD titles. Customers found out, compared notes, and the backlash was immediate. Amazon apologized, said the discounts had been assigned at random and never on the basis of demographics, and refunded 6,896 customers who had paid more than others during the test period.
Three separate risks are worth keeping apart, because they need different responses:
- Trust. Two people who paid different amounts for the same item on the same day will find each other. Household members share a screen, shoppers compare in a group chat, and a price screenshot spreads faster than any apology.
- Consumer-protection rules. Requirements differ by market. The EU's Omnibus Directive obliges traders to disclose a price that has been personalized by automated decision-making. Other jurisdictions have their own rules on price display, advertised prices, and reference prices.
- Discrimination. Assigning prices with any input that correlates with a protected characteristic (location can be one) is a different and more serious problem than random assignment. Amazon's public defense in 2000 rested precisely on the randomness of its assignment.
None of this makes pricing untestable. It makes the naive design (live, simultaneous, different numbers, all traffic) the wrong default. Get the design reviewed before you write the hypothesis, not after the test ends.
Safer ways to test pricing
Rank your options by how much exposure they create, and take the cheapest one that still answers your question.
| Test design | What actually changes | Primary metric | Trust exposure |
|---|---|---|---|
| Price framing and presentation | Monthly vs annual, per-unit vs total, plan order, anchoring | Revenue per visitor | None. Every visitor can buy at the same price. |
| Bundles and offer structure | What is included at a price, not the price of a single item | Revenue and margin per visitor | Low. Bundles are visibly different products. |
| Free-shipping or discount threshold | The order value at which shipping becomes free | Revenue per visitor and AOV | Low. Published as a public offer rule. |
| Sequential price test | The price itself, one period at a time | Revenue per visitor, period vs period | Low. No two shoppers see different prices at once. |
| Live simultaneous price test | The price itself, two numbers at the same moment | Revenue and margin per visitor | High. Legal review required before launch. |
Two further guardrails apply to any test that does move the number:
- Test on new visitors only. Existing customers have a reference price in their heads and, often, an old confirmation email. New-visitor targeting removes most of the comparison risk and gives you a cleaner read on acquisition anyway.
- Honor the lower price. If a variant loses, the people who saw the higher price should not be the ones who pay for the experiment. Decide the make-good policy before launch, not during the backlash.
A wider set of pricing levers, including the demand-driven approach that airlines and marketplaces use, is covered in our guide to dynamic pricing in eCommerce. The pricing strategies you might put into a test (base price adjustments, discounts, bundling) are laid out in this pricing strategy handbook.
Why revenue per visitor is the metric, not conversion rate
This is the single most common way a pricing test goes wrong. Conversion rate is the default success metric in most testing tools, and for a pricing test it is close to useless on its own. Drop the price far enough and conversion rate goes up every time. That is not an insight, it is arithmetic.
Work through a small example. The control price is $50. Out of 10,000 visitors, 300 buy, which is a 3% conversion rate and $15,000 in revenue, so RPV is $1.50. The variant price is $60. Out of 10,000 visitors, 270 buy, a 2.7% conversion rate and $16,200 in revenue, so RPV is $1.62. Conversion rate says the variant lost. Revenue says it won by 8%.
Then add the metrics that decide whether the win is real:
- Gross margin per visitor. Discount variants can beat the control on revenue and still lose on margin. If unit costs differ between variants, margin is the honest primary metric.
- Average order value. Tells you whether a bundle or threshold test moved basket composition or just moved who bought.
- Refunds and returns. A price that pushes hesitant buyers over the line often shows up as returns 30 days later.
- Repeat purchase rate and customer lifetime value. Discount-acquired customers repeat less. A test that wins on first-order revenue and loses on second orders has cost you money.
- Churn, for subscription businesses, where the price change compounds every month.
Set the revenue goal before launch, in the tool. Deciding after the fact which metric you meant is how teams talk themselves into a result.
Set revenue per visitor as the primary goal, target new visitors only, and split traffic at random.
See how Omniconvert Explore runs it →How to design a pricing split test
-
Write a specific, falsifiable hypothesisNot "test a higher price" but "raising the enterprise plan from $200 to $220 will increase revenue per visitor by at least 5% because our win/loss notes show price is rarely the stated objection." A hypothesis names the change, the expected effect, the metric, and the reason.
-
Choose the least exposed design that answers the questionWork down the table above. If a framing or bundle test answers the question, run that. If the number itself has to move, prefer sequential or new-visitors-only, and send the design to legal before you build it.
-
Decide who the test may touchExclude existing customers, recent purchasers, and anyone holding a quote or an active cart. Decide the geographies. Every exclusion is also a decision about which population your result generalizes to.
-
Set revenue per visitor as the primary metricOne primary metric, decided in advance. Add margin per visitor, AOV, refunds, and repeat rate as guardrails. Guardrails do not declare winners; they veto them.
-
Calculate sample size and run length before launchDecide the smallest effect worth acting on, then work out how many visitors per variant that requires. Because order values vary, revenue per visitor is noisier than conversion rate and needs more traffic. Then commit to a fixed end date covering at least one full business cycle.
-
Run it to completion, then look past the test windowNo stopping the moment a variant looks ahead. When it ends, check the guardrails and then check the cohort 30 to 90 days later for refunds, repeat orders, and churn.
If you want to test several elements of the pricing page at once (the plan order, the headline, the guarantee copy), multivariate testing handles combinations, at the cost of needing considerably more traffic.
Segmenting the audience
Not everyone reacts to a price the same way, so a single blended number can hide two opposite effects that cancel out. The segments worth defining before launch:
- Geography. Buying power, expectations, and consumer-protection rules all vary by market. Many teams keep the base price constant across regions and test the offer around it (shipping thresholds, promotional bundles) instead.
- New versus returning. The most important split in any price test. Returning customers hold a reference price; new visitors do not. An introductory offer that works on new buyers can insult loyal ones.
- Acquisition channel. Discount-led paid traffic is more price-sensitive than organic or direct traffic. A price that wins on one channel can lose on another.
- Buying behavior. Bulk and frequent buyers respond to volume tiers and loyalty rewards, not to a change in the unit price. Testing tiers on them, and framing on everyone else, often beats a single blended test.
Segment analysis after the fact is fine for generating the next hypothesis. It is not fine for declaring a winner: slice a flat result enough ways and one slice will always look significant.
Testing prices in a physical store
A grocery chain testing a cereal price lowers it in one set of stores and holds it in another, then compares units sold and category revenue. The logic is identical to an online split test. The execution is harder, because you cannot randomize individual shoppers, only locations.
Three practical rules. Match the store groups on size, location type, and past sales for the same weeks last year before you start, or you are measuring the stores rather than the price. Run long enough to average out weather, paydays, and local events, which usually means several weeks rather than several days. And do not compare a store running a promotion to one that is not, unless the promotion is the thing you are testing.
What makes a pricing result trustworthy
Statistical significance is the standard test of whether a difference is likely to be real rather than random variation. It is necessary and it is not sufficient, and pricing is where the gap between the two hurts most, because a wrong price ships to every order.
Four things separate a result you can act on from a number on a dashboard:
- Adequate sample size, calculated before launch. Work out how many visitors per variant you need to detect the smallest effect you would act on. Doing this afterwards, when you already know the outcome, is not a calculation. A free tool such as Optimizely's sample size calculator does the arithmetic. Our guide to statistical significance covers the mechanics.
- A full business cycle. Weekday and weekend buyers differ. Payday weeks differ. Two to four complete weeks is a reasonable floor for most eCommerce brands.
- No peeking and stopping. Checking daily and stopping the moment a variant leads is the fastest way to manufacture a false winner, and it is especially easy with revenue metrics, where one large order can swing a chart.
- Practical significance. A statistically significant 0.4% revenue lift that requires a repricing project across the catalog is not a win. Decide the threshold worth acting on when you write the hypothesis.
Do not forget how the price is presented
Pricing is not only the number. Framing a price as a monthly fee rather than an annual one changes how affordable it feels, even when the total is identical. Anchoring a $49 plan next to a $99 plan changes what $49 means. A page that explains what the buyer gets for the money can lift revenue without touching the price at all.
This is the most under-used part of pricing optimization, and it is the part with no trust risk whatsoever. Run tests on the layout and copy of the pricing page, then ask buyers directly with an on-site survey what nearly stopped them. The answers usually make the next price hypothesis a much better one than the price you were guessing at.
Frequently Asked Questions
Split testing for pricing is an experiment that shows different price points, offers, or price presentations to comparable groups of visitors, then compares the results to find which one produces the most revenue. One group sees the current price (the control), the other groups see a variant. The comparison is only valid if the groups are split at random and the test runs long enough to collect a reliable sample.
It depends on the market and on how the price is set. In the EU, the Omnibus Directive (EU) 2019/2161 requires traders to tell consumers when a price has been personalized on the basis of automated decision-making. Pricing that differs by a protected characteristic raises discrimination questions in many jurisdictions. This article is not legal advice. Before you run a live price test, get the design reviewed by your own legal counsel.
Revenue per visitor is the right primary metric for a pricing test, not conversion rate. Price moves conversion rate and order value in opposite directions, so a cheaper variant almost always wins on conversion rate while losing money. Revenue per visitor multiplies the two together. Track gross margin per visitor as well when the variants have different unit costs, and watch refunds and churn afterwards.
You can test price presentation and framing (monthly versus annual, per-unit versus total), the order and anchoring of plans, bundles, free-shipping thresholds, payment options, guarantees and returns messaging, and the copy that explains the value behind the price. These change what the price means to the buyer without charging two people different amounts for the same item at the same moment.
Run a pricing test for at least one full business cycle, which for most eCommerce brands means two to four complete weeks, and until it reaches the sample size your power calculation asked for. Revenue per visitor is a noisier metric than conversion rate because order values vary, so pricing tests usually need more traffic than layout tests. Never stop the moment a variant looks like it is winning.
Sequential price testing means running one price for a set period, then switching to the other price for an equal period, instead of showing both prices at the same time. Nobody can compare two prices side by side, so the trust risk largely disappears. The cost is precision: seasonality, promotions, and traffic mix change between the periods, so you must match the periods carefully and treat the result as directional.
Segment by geography, by new versus returning customer, by acquisition channel, and by buying behavior such as bulk or frequent purchasing. Segmentation serves two purposes: it keeps a test off audiences it should not touch, such as existing customers who already bought at the old price, and it shows which groups actually drove a result that the blended average hides.
Yes, usually by running one price in a set of stores and the current price in a matched set of comparable stores over the same period. Store-level tests are noisier than online tests because footfall depends on weather, local events, and time of day, so you need several stores per arm and a longer run. Match the store groups on size, location type, and past sales before you start.
Do not open with a live price test. Start with the questions that carry no trust risk: how the price is framed, what the bundle contains, where the free-shipping threshold sits, and how well the page explains the value behind the number. Those tests answer most pricing questions and they can run today. When you do need to move the number itself, pick the safest design that still answers the question, restrict it to new visitors or run it sequentially, set revenue per visitor as the goal, size the test before you launch it, and have counsel look at the design first. Then keep watching after the test ends. A price that wins on week-one revenue and loses on repeat purchases has not won anything.
Run your pricing experiments in Omniconvert Explore
Omniconvert Explore runs A/B and multivariate tests with random assignment, audience targeting for new-visitor-only or geography-based tests, revenue goals, and on-site surveys to ask buyers what the price meant to them. 70,000+ experiments across 7,000+ websites, with a 23.2% average uplift.