WoluTools
← About this tool

analytics · Browser tool

A/B Test Calculator

Size a test before you run it, or check a finished one for statistical significance with a z-test on two proportions.

Runs in your browserNothing is uploadedHow it works
Free toolNo account · no job limitFree, unlimitedRuns offline once the page has loaded.

What do you need?

Mode

Test parameters

Effect is measured as
Test direction

How many visitors you need

Sample size grows roughly with the square of the effect you want to catch. Halving the minimum detectable effect multiplies the traffic you need by about four, which is usually the reason a test is not worth running.

What this does

Both calculations treat a conversion rate as a binomial proportion and use the normal approximation to its sampling distribution. Sizing solves the standard two-proportion power equation for n; the significance check runs a z-test on the difference between two proportions, pooling the rates under the null hypothesis for the test statistic and keeping them separate for the confidence interval. The approximation is reliable once both groups expect at least about ten conversions and ten non-conversions. Below that, treat the output as a rough guide and use an exact test.

Where the normal curve comes from

There is no closed form for the normal cumulative distribution, so this page computes it from a rational approximation of the error function — the Abramowitz and Stegun 7.1.26 form, good to about 1.5×10−7. The inverse direction, turning a probability such as 0.975 back into a z-score, uses Peter Acklam's rational approximation, accurate to roughly 1.2×10−9. Both are far tighter than the uncertainty in any real conversion data, but they are approximations and the last digit of a p-value near zero should not be read too literally.

What a p-value is not

A p-value of 0.03 says that if the two variants really performed identically, a gap at least this large would turn up about 3% of the time. It does not say there is a 97% chance the variant is better, and it says nothing at all about how much better. Those are different questions and the second one is answered by the confidence interval, not the p-value.

Stopping early breaks the maths

The sample size you calculate is a commitment, not a target to watch creep upwards. If you check the result every morning and stop the moment it crosses 0.05, the real false-positive rate climbs well past the 5% you asked for — with daily peeking over a few weeks it can reach a quarter or more. Fixed-horizon tests like this one assume you look once, at the end. Sequential designs exist for tests you want to monitor, and they need a different calculation.

Statistical and practical significance

A test with enough traffic will eventually declare a 0.05% improvement significant. That is a true finding and usually a worthless one. Decide beforehand how small a lift would still be worth the engineering cost, and read the bottom end of the confidence interval rather than the headline number — if the interval runs from +0.1% to +3%, the honest answer is that you do not yet know whether the change is worth shipping.