/CRO and speed

Shopify A/B testing: do you have enough traffic?

August 2, 2026 · WOCX

Most Shopify stores do not have enough traffic to run a meaningful A/B test, and the arithmetic is not close. To detect a 20% improvement on a store converting at 2%, you need roughly 39,200 visitors split across the two variants before the result means anything. At 100 visitors a day that is 392 days. The test would finish more than a year after it started, by which point the season, the ad account and probably the product range have all changed.

That is an unpopular thing for an agency to publish, because A/B testing sells well. It is also the single most useful number a small store owner can have, because it converts a vague ambition into a yes or no. This piece gives you the numbers for your own traffic, explains why the thresholds are so high, and sets out what to do when the answer is no.

Why the numbers are so large

Conversion rate is a rare event, and rare events need large samples to measure reliably. At a 2% conversion rate, 98 out of every 100 visitors do nothing, so the signal you are trying to detect lives in a very thin slice of the traffic. Random variation in that slice is large relative to the effect you are looking for. If 200 visitors produce 4 orders one week and 6 the next, that is a 50% swing that means nothing at all, and any test short enough to be convenient will be dominated by exactly that kind of noise.

The standard requirement is a result you would only see by chance 5% of the time, with an 80% chance of catching a real effect if one exists. Those two thresholds are what generate the sample sizes below. You can lower them, and plenty of testing tools quietly do, but a test run at looser thresholds is not a cheaper test. It is a test that will tell you things that are not true, and acting on those is worse than not testing at all, because you will ship the change and then defend it.

The numbers for your store

Sample sizes below are per variant and total, using the standard approximation for 95% confidence and 80% power. Find the row closest to your own conversion rate, then read across to the smallest lift you would realistically ship a change for. That total is your entry ticket: below it, the test cannot distinguish your change from ordinary weekly variation, no matter how carefully it is built or how confident the tool’s dashboard looks. Note that the numbers describe visitors to the tested page rather than to the whole store, which matters because a test on a single product page draws from a fraction of site traffic. If that product page sees a fifth of your visitors, multiply the days accordingly. Most merchants who believe they have enough traffic have compared the requirement against sitewide sessions rather than against the page actually under test.

Your conversion rate Lift you want to detect Visitors per variant Total visitors needed
1.4% 10% 112,686 225,372
1.4% 20% 28,171 56,342
1.4% 50% 4,507 9,014
2.0% 10% 78,400 156,800
2.0% 20% 19,600 39,200
2.0% 50% 3,136 6,272
3.0% 20% 12,933 25,866
3.0% 50% 2,069 4,138

Two patterns matter more than any single row. First, the requirement scales with the square of the effect: halving the lift you want to detect quadruples the traffic you need. That is why chasing small wins is out of reach for almost everyone, and why the honest small-store strategy is to make changes large enough to be obvious. Second, a higher baseline conversion rate makes testing cheaper, so the stores best equipped to test are the ones already converting well.

Against the wider picture, the average Shopify store converts at about 1.4%, the top 20% at 3.1% to 3.5%, and the top 10% at 4.7% to 5.2%. A typical store is therefore sitting on the hardest row in that table.

How long that takes in real days

Daily visitors Days to detect a 20% lift at 2% baseline
100 392 days
500 79 days
1,000 40 days
5,000 8 days

Read that table before installing a testing tool, because it decides the answer for you. Under a few hundred visitors a day, valid testing of ordinary changes is not available, and no tool fixes that. Between roughly 500 and 1,000 a day you can test a small number of substantial changes per quarter, which means test selection matters more than test execution. Above a few thousand a day, testing becomes a genuine operating rhythm and the usual advice applies.

There is one more constraint people forget. A test needs to run for full weeks regardless of traffic, because buying behaviour differs by day, and a test that starts on a Monday and ends on a Thursday has sampled an unrepresentative week. If your maths says 9 days, run 14.

DAYS TO A VALID RESULT (20% LIFT, 2% BASELINE) 100/day 392 500/day 79 1,000/day 40 5,000/day 8
The same test, four stores. Traffic decides whether A/B testing is a tool available to you at all.

What to do when the answer is no

Not being able to test is not the same as not being able to improve, and the alternative is not guessing. It is a different method with different evidence.

Ship changes in deliberate bundles rather than one at a time. If you cannot isolate a variable, stop trying, and instead group several changes that share a rationale so the combined effect is large enough to see in the ordinary numbers. This is exactly what we did on the buy box work, where five changes around the add-to-cart button moved a product page from 3.9% to 5.3%. We could not attribute that to any single change and we do not claim to. The bundle was the unit of work because isolated testing was never viable at that traffic level, and that is the honest trade rather than a shortcut.

Use evidence that does not need statistical power. Session recordings and heatmaps show you where people hesitate without needing a sample size. Your own search logs tell you what visitors could not find. Support tickets and pre-sale questions name the objections your product page failed to answer. None of this proves causation, and none of it needs 39,200 visitors either.

Fix the things that are wrong regardless of test results. A checkout that switches currency, an add-to-cart button below the fold on mobile, a page taking four seconds to load: these do not need testing, because there is no version of the store where they are the better option. When we measured 7 real DTC product pages at a phone viewport, add-to-cart sat below the fold on 4 of 6. Nobody needs a test to know that is worth fixing.

Your traffic The honest method
Under 500/day Bundled changes, qualitative evidence, fix obvious defects
500 to 1,000/day 2 to 4 substantial tests a year, chosen carefully
Over 5,000/day Continuous testing, the standard playbook applies

The mistakes that produce false confidence

Three habits produce results that feel solid and are not, and all three are more common than genuine testing errors because they feel like diligence rather than mistakes. Each one converts an underpowered test into a confident decision, which is worse than having run no test at all: a store with no data knows it is guessing, while a store with a bad result believes it has evidence and will defend the change for months. The most common one is stopping a test the moment it looks like it is winning. Conversion data wanders, an early lead is usually noise, and a test called at day 3 because the variant was up 30% is not a result. Decide the sample size before you start and do not look at the outcome until you reach it.

The second is running many tests at once on a small store and treating the winners as real. Run 20 underpowered tests and a few will show a strong effect by chance alone. That is not a testing programme, it is a random number generator with a dashboard.

The third is testing tiny changes because they are easy to build. Button colour is the classic. At a 2% baseline, a genuine 2% relative improvement from a colour change would need well over a million visitors to detect, so what you are really measuring is noise. If a change is not big enough to plausibly move the number by a fifth, it is not a candidate for testing at your traffic. The statistical significance framing is worth understanding for exactly this reason.

FAQ

How much traffic do I need to A/B test?

To detect a 20% lift at a 2% conversion rate, about 39,200 visitors total. At 100 visitors a day that is 392 days, which means valid testing is not available to most small stores.

Can I test with less traffic?

Only for very large effects. Detecting a 50% lift at 2% needs about 6,272 visitors rather than 39,200, so testing is possible if the change is dramatic.

Why do smaller lifts need more traffic?

The requirement scales with the square of the effect. Halving the lift you want to detect quadruples the visitors needed, which is why small optimisations are out of reach at low traffic.

How long should a test run?

Full weeks, always, even if the maths says fewer days. Buying behaviour varies by weekday, so a part-week test samples an unrepresentative slice.

What should I do instead of testing?

Ship deliberate bundles of related changes, use session recordings and search logs for evidence that needs no sample size, and fix clear defects without waiting for proof.

Not sure whether your store has the traffic to test, or what to change first if it does not? Send us the store and we will tell you honestly which situation you are in. Free, no obligation, usually a reply within the hour.