analytics · Browser tool
A/B Test Calculator
Size a test before you run it, or check a finished one for statistical significance with a z-test on two proportions.
What do you need?
Test parameters
How many visitors you need
Sample size grows roughly with the square of the effect you want to catch. Halving the minimum detectable effect multiplies the traffic you need by about four, which is usually the reason a test is not worth running.
Observed results
Pick the direction before you look at the numbers. Switching to a one-sided test after seeing which way the result went halves your p-value for free and is not a valid move.
What the numbers say
What this does
Both calculations treat a conversion rate as a binomial proportion and use the normal approximation to its sampling distribution. Sizing solves the standard two-proportion power equation for n; the significance check runs a z-test on the difference between two proportions, pooling the rates under the null hypothesis for the test statistic and keeping them separate for the confidence interval. The approximation is reliable once both groups expect at least about ten conversions and ten non-conversions. Below that, treat the output as a rough guide and use an exact test.
Where the normal curve comes from
There is no closed form for the normal cumulative distribution, so this page computes it from a rational approximation of the error function — the Abramowitz and Stegun 7.1.26 form, good to about 1.5×10−7. The inverse direction, turning a probability such as 0.975 back into a z-score, uses Peter Acklam's rational approximation, accurate to roughly 1.2×10−9. Both are far tighter than the uncertainty in any real conversion data, but they are approximations and the last digit of a p-value near zero should not be read too literally.
What a p-value is not
A p-value of 0.03 says that if the two variants really performed identically, a gap at least this large would turn up about 3% of the time. It does not say there is a 97% chance the variant is better, and it says nothing at all about how much better. Those are different questions and the second one is answered by the confidence interval, not the p-value.
Stopping early breaks the maths
The sample size you calculate is a commitment, not a target to watch creep upwards. If you check the result every morning and stop the moment it crosses 0.05, the real false-positive rate climbs well past the 5% you asked for — with daily peeking over a few weeks it can reach a quarter or more. Fixed-horizon tests like this one assume you look once, at the end. Sequential designs exist for tests you want to monitor, and they need a different calculation.
Statistical and practical significance
A test with enough traffic will eventually declare a 0.05% improvement significant. That is a true finding and usually a worthless one. Decide beforehand how small a lift would still be worth the engineering cost, and read the bottom end of the confidence interval rather than the headline number — if the interval runs from +0.1% to +3%, the honest answer is that you do not yet know whether the change is worth shipping.