Meelu

Free A/B test calculator

You ran a test and one version is ahead. Is that a real win or just noise? Enter the visitors and conversions for each version and find out — including whether it is safe to stop yet. Everything runs in your browser. No sign-up, nothing uploaded.

Visitors and conversions per variant
Confidence
Hypothesis
Multiple comparisons
Bayesian prior

Every conversion rate between 0 and 100% equally plausible before the data. The honest default when you genuinely know nothing, and what almost every Bayesian calculator uses.

Every calculation runs in this browser tab. Nothing is uploaded, nothing is stored.

Variant B beats Control

Variant B converted at 4.56% against 4.00%, a relative lift of 14.0%. p = 0.0287, under the 0.0500 threshold. The Bayesian view agrees: 98.6% probability it is genuinely better, and shipping it costs you 0.001pp of conversion rate on average if it is not. The honest range for the true lift is 1.4% to 28.2% — plan on the low end.

Relative lift+14.0%4.00% → 4.56%
p-value (threshold 0.0500)0.0287z = 2.187
P(variant is better)98.6%Bayesian, closed form
Expected loss if you ship it0.001ppaverage conversion rate given up if it is actually worse

Every arm

25,000 visitors · 1,070 conversions
ArmVisitorsConversionsRate95% intervalLiftP(best)
Controlcontrol12,5005004.00%3.67%4.36%1.4%
Variant Bsignificant12,5005704.56%4.21%4.94%+14.0%98.6%

Per-arm intervals are Wilson score intervals, which behave sensibly at small counts where the textbook Wald interval can run past 0% or 100%. P(best) comes from 40,000 draws from the posteriors.

Posterior distributions

3.13%3.72%4.32%4.91%5.50%density
Control · Beta(501.0, 12001.0)Variant B · Beta(571.0, 11931.0)

Each curve is the range of conversion rates still plausible for that arm after the data. Overlap is the whole story: if two curves sit on top of each other, no amount of significance testing will tell them apart. Prior: Uniform — Beta(1, 1) — Beta(1.00, 1.00).

Variant B vs Control

98.6% to be better

Frequentist

z statistic (pooled variance)
2.1873
Chi-square, 1 df
4.7842 (same test, same p)
p-value, two-sided
0.0287
p-value, one-sided
0.0144
Fisher exact, two-sided
0.0310
95% CI, absolute lift
+0.06pp to +1.06pp
95% CI, relative lift
+1.4% to +28.2%
Smallest lift this sample could catch
+17.4% at 80% power

Bayesian

P(variant beats control)
98.56% · closed form
P(variant is the best of all arms)
98.58%
Expected loss if you ship the variant
0.0013pp
Expected loss if you keep the control
0.5612pp
95% credible interval, absolute lift
+0.06pp to +1.06pp
95% credible interval, relative lift
+1.3% to +28.1%
Posterior
Beta(571.0, 11931.0)

What the p-value is not. p = 0.0287 does not mean there is a 97.1% chance the variant is better. It is the probability of seeing a gap this big or bigger if the two versions were identical — a statement about the data under an assumption, not about the assumption given the data. The number that answers the question you are actually asking is the Bayesian one above it, and the two frequently disagree in tone even when they agree on the arithmetic.

95% confidence” is not a 95% chance of winning. It means that if you repeated this experiment forever, 95% of the intervals produced would contain the true lift. This particular interval either contains it or it does not. And a lift can be perfectly real and still not worth having: yours could be as low as +1.4% — budget for that number, not the headline one.

  • No account
  • Nothing uploaded
  • Tells you when to stop
  • Warns if the test is broken
  • Up to 6 versions

What you get back

Most free calculators answer one question: is the difference significant? That is the easy part. The hard parts are knowing when you are allowed to stop, spotting that the experiment itself went wrong, and not reading more into a result than is there.

A straight answer on the winner

Did the variant beat the control, or is the gap just noise? You get the verdict in a sentence, with the numbers behind it if you want them.

How likely the variant is better

A plain probability that the variant really is the better version, and what it would cost you on average if you picked the wrong one.

Safe to check any day

Watching a test daily and stopping the moment it looks good is how teams ship changes that do nothing. This gives you a reading you can act on even if you have been peeking.

Warns you when the test is broken

If traffic did not split the way you set it up, the result is not trustworthy no matter what it says. You get told, instead of being left to find out later.

How many visitors you need

Work out the traffic needed to spot the size of lift you care about, and how long that will take at your current weekly visitors.

Reach an answer with less traffic

If you have data on the same visitors from before the test, it can be used to cut through some of the noise and shorten the test.

How to check A/B test significance

  1. Step 1

    Enter visitors and conversions

    One row per version you tested, up to 6. The first row is your control. If you did not split traffic evenly, say what the split was meant to be.

  2. Step 2

    Read the verdict, then the range

    The plain answer comes first. Then look at the range of possible lift: a headline of +14% that could really be anything from +1% to +29% is not a 14% win.

  3. Step 3

    Check before you stop

    If you have been watching the test day by day, use the reading that accounts for that. It will usually tell you to keep running a little longer.

The four things people get wrong

These are the mistakes that turn a testing programme into a coin toss, in rough order of how much damage they do.

Stopping the moment it looks like a win
We ran 20,000 fake tests where both versions were identical, checking each one daily for two weeks. Stopping the first time the numbers looked good declared a winner in more than one in five of them — every one of those a false win. If you are going to check daily, use the reading on this page that is built for it.
Thinking 95% confident means 95% sure the variant wins
It does not. The everyday significance number tells you how surprising your gap would be if the two versions were really the same. If you want the odds that the variant is genuinely better, this page gives you that separately — and it is a different number.
Quoting the headline lift and ignoring the range
A +14% lift that could realistically be anything from +1% to +29% is not a 14% lift. The range is the honest answer. The wider it is, the less you know.
Shipping a win that is worth nothing
With enough traffic, even a 0.2% lift nobody will ever notice will register as a real difference. Decide before you start how big a change is worth shipping, and judge the result against that.

What it does not do

  • Conversion rates only. Revenue per visitor, order value and time on site are a different kind of measurement and the winner check here will not handle them.
  • Fixed traffic splits only. If your testing platform shifts traffic towards whichever version is winning as the test runs, none of these answers apply.
  • It only sees totals. A lift that shows up in week one and fades by week three is invisible here. Run whole weeks and look at the trend yourself.
  • It cannot tell you why. You get whether the numbers differ and by how much. The reason is your job — occasionally it turns out to be “the variant was broken on Safari”.
  • Nothing is saved. No account, no history, no monitoring. Copy the summary if you want to keep it.

Frequently asked questions

What is A/B testing?

A/B testing is running two versions of a page, email or feature at the same time and splitting traffic between them to see which one performs better. Version A is usually the current design and version B is the change you want to try. Because both versions run over the same period with the same kind of visitors, any difference in conversion rate is down to the change rather than the day of the week or a campaign that happened to launch. It is the simplest way to find out whether an idea actually works instead of arguing about it.

What is statistical significance in an A/B test?

Statistical significance means the difference you measured is unlikely to have happened by chance alone if the two versions really performed the same. Conversion rates wobble from day to day, so a small gap between A and B proves nothing on its own. A significance test weighs the size of the gap against how much random variation you would expect at your sample size. Significant does not mean important: a result can be significant and still too small to be worth shipping.

What is a good p-value for an A/B test?

0.05 is the usual threshold, meaning there is a 5% chance of seeing a difference this large if the two versions were identical. Lower thresholds such as 0.01 make you more certain but need more traffic. The p-value is not the probability that your variant is better, and 1 minus the p-value is not your confidence in the winner. Pick the threshold before you start the test, not after you have seen the numbers.

How many visitors do I need for an A/B test?

It depends on your current conversion rate and the smallest lift you care about detecting. As a rough guide, a site converting at 3% that wants to detect a 10% relative lift needs somewhere around 50,000 visitors per variant. Smaller effects need dramatically more traffic: halving the effect you want to catch roughly quadruples the sample. Work the number out with a sample size calculator before you launch, because a test that can never reach significance is wasted traffic.

How long should you run an A/B test?

Run it for at least one full week, and ideally two, even if you hit your sample size sooner. Behaviour differs between weekdays and weekends and between paydays and the rest of the month, so a test that only covers a Tuesday to Thursday measures a narrow slice of your audience. Decide the end date and the sample size up front, then leave the test alone until you reach both. Stopping the moment a result looks good is the most common way teams ship changes that do nothing.

What is statistical power and why does it matter?

Power is the chance your test detects a real effect when one exists, and 80% is the usual target. A test with low power will often come back inconclusive even when the variant genuinely is better, so you throw away a good idea. Power rises with sample size and with the size of the effect you are looking for. Low power is why most tests on low-traffic sites read as flat.

What is the minimum detectable effect?

The minimum detectable effect, or MDE, is the smallest change in conversion rate your test can reliably spot at the traffic you have. If your MDE is 10% and your variant delivers a genuine 3% lift, the test will almost certainly report nothing. Setting a realistic MDE up front keeps you from running tests that were never capable of answering the question. It also forces an honest conversation about whether the change is big enough to be worth testing at all.

Why is peeking at A/B test results a problem?

Checking a running test repeatedly and stopping as soon as it crosses significance inflates your false positive rate, often to 20% or 30% rather than the 5% you think you are accepting. Every extra look is another chance for random noise to cross the line. Either fix the sample size and end date in advance, or use a sequential testing method built to be checked continuously. Peeking with a standard z-test is how teams end up with a pile of wins that never show up in revenue.

A significant test is one number in a bigger picture

A winning variant tells you which page converts better. It does not tell you which channel brought the traffic, or what that traffic was worth. For the first, the attribution model comparison tool shows how differently first-touch, last-touch and time-decay credit the same conversions. For the second, the marketing mix model works out what your spend is actually doing at the channel level. And if the traffic is not arriving in the first place, start with the free SEO audit.

All of them are pieces of Meelu, a desktop app where an AI marketing agent runs your marketing on your own machine — reading your real experiment data instead of numbers you retyped into a form, and telling you when a test is not ready to call. Join the waitlist.