Calculators

A/B test significance and sample size, in one place

Find out whether a result is real or noise, and how many visitors each variant needs before you call a winner.

Significance check

Two-proportion z-test, two-sided. Enter visitors and conversions for each variant.

Variant A (control)
Variant B (challenger)
Enter visitors and conversions for both variants.

Sample size planner

Visitors needed per variant before the test starts. Relative MDE means a 20% change on a 10% baseline is 10% to 12%.

Enter a baseline rate and a minimum detectable effect.

Runs entirely in your browser. Nothing is uploaded.

Method

How it works

Significance uses a pooled two-proportion z-test. Both variants share one estimate of the conversion rate, then the standard error is sqrt(p (1 - p) (1/nA + 1/nB)). The z-score is the difference in rates divided by that standard error. The two-sided p-value is the probability of a z-score at least this extreme in either direction.

z = (pB - pA) / sqrt(p (1 - p) (1/nA + 1/nB)), with p = (cA + cB) / (nA + nB)

Sample size uses the standard two-sample formula for proportions. It needs the baseline rate, the target rate implied by the relative effect, the critical value for your confidence level, and the power term for your chosen power.

n = (z(1-a/2) sqrt(2 pbar (1 - pbar)) + z(1-b) sqrt(p1 (1 - p1) + p2 (1 - p2)))^2 / (p2 - p1)^2

Results are rounded up to whole visitors. Days estimate assumes the same daily traffic is split evenly between the two variants, and that the traffic is representative of a normal week.

FAQ

Questions people ask

What confidence level should I use?

95% is the common default for marketing tests. Use 90% when a wrong call is cheap and speed matters more, and 99% when a wrong winner would be expensive to roll out, such as pricing or checkout changes.

What does the p-value mean?

The p-value is the chance of seeing a gap at least this large if A and B actually convert at the same rate. A p-value of 0.035 means a gap this big would appear by chance about 3.5% of the time. It does not tell you the probability that B is better.

Why is my result not significant if B looks better?

A visible uplift can still be noise when the sample is small. Check the conversions per variant, then keep the test running until the sample size planner says you have enough visitors for the effect you care about.

What is minimum detectable effect?

It is the smallest relative change you want the test to be able to find. A 20% relative effect on a 10% baseline means you are looking for a move from 10% to 12%. Smaller effects need much bigger samples, so pick the smallest change that would still matter to the business.

Should I stop the test as soon as it is significant?

No. Checking repeatedly and stopping at the first significant result inflates false winners. Fix the sample size in advance from the planner, run the test to that size, then read the result once.

Does this work for revenue or average order value?

This calculator is for conversion rates, where each visitor either converts or does not. Revenue per visitor needs a different test, usually a t-test on the per-visitor values.

Next step

Want this done properly across your whole stack?

Tracking, search, automation and reporting, engineered and operated by subimpact network. Start with a free audit.

Get Free Audit