A/B Test Significance Calculator

Find out whether your winning variant is a real improvement or just noise. Get the confidence level, the p-value, and the sample size you'd actually need to call it.

Enter your test results

Variant A (control)
Variant B (challenger)
Confidence level
97.2%
Significant at 95%

The difference is very unlikely to be random chance. Safe to call a winner at the 95% standard.

A conversion rate5.00%
B conversion rate6.00%
Relative uplift+20.0%
Z-score2.19
P-value0.0283
For this effect size you'd need about 8,158 visitors per variant for a properly powered test (80% power, 95% confidence).

Why tests stall, and how to finish them faster

Most A/B tests don't fail because the idea was bad. They fail because they never collect enough conversions to prove anything.

Significance is driven by conversions, not just traffic

The statistics above run on conversions, not raw visitors. A page converting at 1% needs roughly ten times the traffic of a page converting at 10% to detect the same relative improvement. The fewer conversions per visitor, the longer a test takes.

That's why teams with low-converting pages watch tests sit at "trending" for weeks: the sample-size estimate in the calculator tells you exactly how far away the finish line is at your current conversion rate.

Speed up every future test

You can't magic up more traffic. You can convert more of the traffic you already have, and every extra conversion shortens every test you'll ever run.

More conversions per visitor = faster significance. ZipTier's AI assistant engages visitors in conversation, answers objections, and converts more of the traffic you're already splitting, so every A/B test you run reaches significance sooner. And every conversation tells you why a variant wins.

See how ZipTier lifts conversion

How to read your result

The p-value drives the verdict. Here's what each band means before you declare a winner.

P-valueConfidenceVerdictWhat to do
Under 0.0199%+Highly significantVery strong evidence. Ship the winner (or kill the loser).
0.01–0.0595–99%Significant at 95%Meets the standard bar for calling a result. Safe to act.
0.05–0.1090–95%TrendingSuggestive, but below the bar. Keep the test running.
Over 0.10Under 90%Not significantThe difference is within the range of noise. Don't ship yet.

These bands use the two-tailed test convention. A "significant" result can go either way, so check the direction of the uplift before you celebrate.

What statistical significance actually means

When variant B beats variant A in a test, there are always two possible explanations: B is genuinely better, or you just got lucky with which visitors landed on which page. Statistical significance is how you tell those apart. The p-value is the probability of seeing a difference this large if the variants truly performed the same. Pure luck, no real effect. A p-value of 0.03 means that if A and B were identical, only 3% of tests would show a gap this big by chance alone.

Rule of thumb: p-value below 0.05 = at least 95% confidence = you can call a winner. Above 0.05, the "difference" you're seeing is still plausibly random noise.

Why 95% is the convention

The 95% threshold (p < 0.05) isn't a law of nature. It's a convention that balances two kinds of mistake. Set the bar lower and you'll ship "winners" that are actually flukes; set it higher and every test takes far longer to conclude. At 95%, roughly one in twenty truly-no-difference tests will still produce a false winner, which most marketing teams accept as a reasonable error rate. For high-stakes changes like pricing or checkout flows, many teams demand 99% instead.

The danger of peeking

The single most common way A/B tests go wrong is stopping the moment significance appears. P-values wobble as data comes in: a test that will end at p = 0.30 may dip below 0.05 several times along the way. If you check daily and stop at the first green light, your real false-positive rate can be two to three times higher than the 5% you think you're running. Decide your sample size up front (use the estimate in the calculator), run to that number, and only then read the result.

One-tailed vs. two-tailed tests

A one-tailed test asks "is B better than A?" and ignores the possibility that B is worse, which makes significance easier to reach but hides losing variants. This calculator uses a two-tailed test: it asks "are A and B different in either direction?" That's the more conservative choice, and it's why the same data can show 95% confidence in a one-tailed tool but only ~90% here. If a variant can hurt you (it always can), trust the two-tailed number.

Frequently asked questions

Ship winners, not noise

ZipTier's AI assistant converts more of the traffic you're already splitting, so your tests hit significance sooner. And every conversation tells you why a variant won.

Try ZipTier FreeNo credit card required