How do I use statistical significance in testing for ecommerce?

Expert answer · sourced from 0 podcast episodes

Short answer

While it's easy to get lost in the jargon, statistical significance is simply about asking, 'Is this result real or just random luck?' It's a risk management tool that measures your confidence that an A/B test winner is truly better before you change your whole site.

TL;DR

While it's easy to get lost in academic definitions, the consensus among experts is that statistical significance is essentially a risk management tool. On Honest Ecommerce, Adam Kitain puts it simply: it helps you understand the probability that your new version is actually better than the old one. It’s a way of asking, “How confident are we that this result is real and not just a product of random chance?” It’s crucial for making sound decisions in conversion rate optimization.

So what does a significance level, like 90% or 95%, really mean? Marty Greif explained on Firing The Man that a 90% confidence level means that nine times out of ten, your winning variant will truly be a winner, but there's still a one in ten chance it won't be. The higher the number, the more confident you can be. A critical point he makes is that significance isn't driven by the number of visitors in your test, but by the number of conversions. Traffic is a necessary ingredient to get conversions, but the conversion count itself is what matters for the math. Without enough conversions, any lift you see could easily be noise.

This naturally leads to the question of how much data is enough. Several hosts give practical rules of thumb. On Ecommerce Coffee Break, Claus Lauter suggests you need at least 100, and ideally up to 500, transactions for each variant of your test to make an educated decision. Similarly, Yi Hung Lin, also on Ecommerce Coffee Break, recommends a baseline of 100 orders or 10,000 sessions per variant. These are not absolute laws but serve as fantastic starting points. If the performance difference between your variations is very small, you will need a much larger sample size to declare a winner with confidence.

Thinking about data requirements shouldn't be an afterthought. On an episode of Honest Ecommerce, Aaron Zagha, citing his background in statistics, emphasizes that this is one of the most important concepts in testing. He advocates for a proactive approach, using tools like sample size calculators before ever launching a test. This planning step helps you determine if you even have enough traffic to reach a statistically valid conclusion in a reasonable amount of time. If a test would need to run for six months to get a clear result, it's probably not the right test for you to be running right now.

Ultimately, hitting a 95% confidence score isn't the entire goal. The business impact is just as important. Adam Kitain connects the statistical output directly to metrics that matter, like a potential 5% uplift in revenue per visitor. A test could produce a statistically significant result, but if the projected lift is only 0.5% and the change is difficult to implement, is it worth it? Conversely, a test that shows a 20% lift with only 85% confidence might be worth rolling out, depending on your appetite for risk and the cost of being wrong. Statistical significance is a powerful input, but it doesn't replace business judgment. It provides the data you need to make an informed decision, balancing the potential upside against the risk that you’re just seeing things in the noise.

Voices that come up across these episodes

Ask your own question

Get a personalized answer pulled from 23,800 ecommerce podcast episodes.

Ask a question →

More answers