A/B Test Significance Calculator
Find out whether your winning variant is a real improvement or just noise. Get the confidence level, the p-value, and the sample size you'd actually need to call it.
Enter your test results
The difference is very unlikely to be random chance. Safe to call a winner at the 95% standard.
Why tests stall, and how to finish them faster
Most A/B tests don't fail because the idea was bad. They fail because they never collect enough conversions to prove anything.
Significance is driven by conversions, not just traffic
The statistics above run on conversions, not raw visitors. A page converting at 1% needs roughly ten times the traffic of a page converting at 10% to detect the same relative improvement. The fewer conversions per visitor, the longer a test takes.
That's why teams with low-converting pages watch tests sit at "trending" for weeks: the sample-size estimate in the calculator tells you exactly how far away the finish line is at your current conversion rate.
Speed up every future test
You can't magic up more traffic. You can convert more of the traffic you already have, and every extra conversion shortens every test you'll ever run.
More conversions per visitor = faster significance. ZipTier's AI assistant engages visitors in conversation, answers objections, and converts more of the traffic you're already splitting, so every A/B test you run reaches significance sooner. And every conversation tells you why a variant wins.
See how ZipTier lifts conversion →How to read your result
The p-value drives the verdict. Here's what each band means before you declare a winner.
| P-value | Confidence | Verdict | What to do |
|---|---|---|---|
| Under 0.01 | 99%+ | Highly significant | Very strong evidence. Ship the winner (or kill the loser). |
| 0.01–0.05 | 95–99% | Significant at 95% | Meets the standard bar for calling a result. Safe to act. |
| 0.05–0.10 | 90–95% | Trending | Suggestive, but below the bar. Keep the test running. |
| Over 0.10 | Under 90% | Not significant | The difference is within the range of noise. Don't ship yet. |
These bands use the two-tailed test convention. A "significant" result can go either way, so check the direction of the uplift before you celebrate.
What statistical significance actually means
When variant B beats variant A in a test, there are always two possible explanations: B is genuinely better, or you just got lucky with which visitors landed on which page. Statistical significance is how you tell those apart. The p-value is the probability of seeing a difference this large if the variants truly performed the same. Pure luck, no real effect. A p-value of 0.03 means that if A and B were identical, only 3% of tests would show a gap this big by chance alone.
Why 95% is the convention
The 95% threshold (p < 0.05) isn't a law of nature. It's a convention that balances two kinds of mistake. Set the bar lower and you'll ship "winners" that are actually flukes; set it higher and every test takes far longer to conclude. At 95%, roughly one in twenty truly-no-difference tests will still produce a false winner, which most marketing teams accept as a reasonable error rate. For high-stakes changes like pricing or checkout flows, many teams demand 99% instead.
The danger of peeking
The single most common way A/B tests go wrong is stopping the moment significance appears. P-values wobble as data comes in: a test that will end at p = 0.30 may dip below 0.05 several times along the way. If you check daily and stop at the first green light, your real false-positive rate can be two to three times higher than the 5% you think you're running. Decide your sample size up front (use the estimate in the calculator), run to that number, and only then read the result.
One-tailed vs. two-tailed tests
A one-tailed test asks "is B better than A?" and ignores the possibility that B is worse, which makes significance easier to reach but hides losing variants. This calculator uses a two-tailed test: it asks "are A and B different in either direction?" That's the more conservative choice, and it's why the same data can show 95% confidence in a one-tailed tool but only ~90% here. If a variant can hurt you (it always can), trust the two-tailed number.
Frequently asked questions
Ship winners, not noise
ZipTier's AI assistant converts more of the traffic you're already splitting, so your tests hit significance sooner. And every conversation tells you why a variant won.
Try ZipTier FreeNo credit card required