Free tool
A/B Test Calculator: Significance, Sample Size and Verdict
This free A/B test calculator (also called a split test calculator) tells you if your test result is real or just luck. It then tells you what to do next: ship, keep running or stop. It also works out how many visitors you need and how long to run your test.
| Visitors | Conversions | Rate | |
|---|---|---|---|
| Version A | 5.00% | ||
| Version B | 5.60% |
Filled in with the example from this page. Replace it with your own numbers.
Test details (optional)Planned 31,234 per version · 14 days
Add these to run every data check.
Verdict
Keep running
No clear winner yet, and you haven't reached your planned visitors.
What to do Keep the test running. Don't stop just because B looks good today.
1 data check needs attention
- Traffic split looks healthy: 10,000 vs 10,000 visitors.
- Stopped too early: 10,000 of your planned 31,234 visitors per version.
- Ran for 14 days.
- At least 100 conversions per version.
Waiting weeks for an answer? We can help you test changes big enough to show up. Book a call
Running tests and not sure you can trust them? We set up experimentation that gives answers you can act on.
Book a callHow do I use this A/B test calculator?
There are two tabs. Use Plan a test before your test starts. Use Check a result when your test is done.
To check a result, enter four numbers:
- visitors who saw version A
- conversions from version A
- visitors who saw version B
- conversions from version B
Then pick how sure you want to be. 95% is the usual choice. The calculator shows you:
Under Test details, you can also add your planned visitors and how many days the test ran. These turn on every data check.
- the conversion rate of each version
- how much better or worse B did (the lift)
- the p-value and the likely range of the real lift
- the chance that B is better than A
- a clear verdict, with the next step to take
To plan a test, enter your current conversion rate, the smallest improvement you care about, and your weekly traffic. The calculator tells you how many visitors each version needs and how many weeks to run the test. Click Use this plan to check the result to carry your planned visitors over to Check a result, so the calculator can tell you if you stopped too early. Or click Copy plan link to save the plan, and open the link when your test is done.
How do I read my A/B test result?
You get one of five verdicts. Each one comes with a next step, so you don't have to work out what the numbers mean on your own.
| Verdict | What it means | What to do |
|---|---|---|
| Ship B | B is better, you had enough visitors, and your data looks healthy | Roll out B. Share the lift as a range, not one number. |
| Keep running | No clear winner yet, or B looks like a winner but you haven't reached your planned visitors | Keep the test running. Don't stop just because B looks good today. |
| No real difference | You reached your planned visitors, and any difference is too small to matter | Stop the test. Keep A, or choose B for other reasons like cost or design. |
| Keep A | B is worse than A, and the gap is unlikely to be luck | Stop the test and keep A. You don't need to wait for your planned visitors when B is clearly losing. |
| Don't trust this | Something is wrong with the data, like an uneven traffic split | Fix your tracking or setup and run the test again. |
Here's why the verdict matters. Say 10,000 people saw each version. A got 500 conversions and B got 560. That's a 12% lift, which sounds like a clear win.
But the p-value is 0.058. That's just above the usual 0.05 cut-off, so the result is not yet proven. If you haven't reached your planned number of visitors, the right call is "Keep running", not "Ship B".
Is my A/B test result statistically significant?
Your result is statistically significant when the p-value is below 0.05, if you're testing at 95% confidence. In simple terms, it means the difference between A and B is unlikely to be random luck.
It does not mean there's a 95% chance that B is better. That's a common mix-up.
What does the p-value actually tell me?
The p-value answers one question: if A and B were actually the same, how often would you see a difference this big just by chance?
A p-value of 0.03 means you'd see a gap this large only about 3 times in 100 if there were no real difference. Rare, so the difference is probably real.
The American Statistical Association's statement on p-values makes two things clear. The p-value is not the chance that your idea is right. And it doesn't tell you how big or important the effect is.
We looked at the 10 top-ranking A/B test calculators on Google. Three of them describe a significant result as "95% sure" or "99% certain" that B is better. That's not what the maths says, and this calculator avoids that wording.
What is a confidence interval in A/B testing?
The lift is one number. The confidence interval is the range where the real lift probably sits. The range tells you much more.
In our example, the lift is +12%. But the range runs from −0.4% to +24%. So the real effect could be nothing at all, or it could be a big win. If you only report "+12%" to your team, you're hiding how uncertain the result is.
The calculator shows both, and draws the range as a bar so you can see when it crosses zero.
What is the probability that B beats A?
It's a second way of reading the same data, called the Bayesian method. It answers a question most people actually ask: how likely is it that B is better than A?
In our example, the chance that B is better is about 97%. That sounds decisive, even though the p-value just missed 0.05.
Both numbers are correct. They just answer different questions. And a 97% chance that B is better still allows B to be only a tiny bit better, which may not be worth the change. Decide before the test which number your team will use to make the call.
How many visitors do I need for an A/B test?
You need enough visitors to spot the smallest improvement you care about. The lower your conversion rate and the smaller the improvement, the more visitors you need.
The table shows visitors needed for each version, at 95% confidence and 80% power, using a two-sided test. Power is the chance your test spots a real improvement when there is one. At 80%, you'd catch it 8 times out of 10.
| Current conversion rate | To spot a 5% lift | To spot a 10% lift | To spot a 20% lift |
|---|---|---|---|
| 1% | 637,000 | 163,100 | 42,700 |
| 2% | 315,200 | 80,700 | 21,100 |
| 5% | 122,100 | 31,200 | 8,200 |
| 10% | 57,800 | 14,800 | 3,800 |
A 10% lift on a 2% conversion rate means going from 2% to 2.2%. Notice the pattern: if you want to spot a lift half as big, you need about four times as many visitors. That's why small improvements are so hard to prove.
What minimum detectable effect should I choose?
Your minimum detectable effect (MDE) is the smallest improvement worth acting on. Pick it based on what matters to your business, not what you hope to see.
Ask yourself: what lift would make this change worth it? If a new checkout design takes two weeks to build, a 1% lift may not be worth it. If only a 10% lift would change your decision, plan the test around 10%.
How do I run an A/B test with low traffic?
If your traffic is low, you have two choices. Accept that you can only spot bigger improvements, or test bigger changes. Small tweaks like button colours rarely move the numbers enough to show up on low traffic.
Here's an example. Your page converts at 3%, and each version gets 10,000 visitors a week. The smallest lift you can reliably spot is:
- about 17% after 2 weeks
- about 12% after 4 weeks
- about 8% after 8 weeks
The Plan a test tab does this math for your own numbers in its "What can you spot with your traffic?" table, so you know if a test is worth running before you build it.
How long should an A/B test run?
Run the test until you reach your planned number of visitors, and for at least one full week. Always run in whole weeks, so both versions see weekday and weekend visitors.
Don't stop the test the moment it looks like a winner. If you check every day and stop at the first good-looking result, you'll pick false winners far more often than you think. Decide the number of visitors up front, and wait for it.
Can I trust my A/B test result?
A result is only useful if the data behind it is healthy. Before giving a verdict, the calculator checks for these problems:
- Uneven traffic split. You planned 50/50 but got 10,400 vs 9,600 visitors. A gap like that is very unlikely by chance. It usually means your setup or tracking is broken. This is called a sample ratio mismatch. The calculator flags a split when a gap that big would happen less than 1 time in 2,000 by chance, the same cut-off Microsoft's experimentation team uses.
- Stopped too early. The test ended before reaching its planned number of visitors.
- Too short. The test ran for less than 7 days, so it missed part of the week.
- Too few conversions. Fewer than about 100 conversions per version makes the result shaky.
If the traffic split is uneven, the verdict becomes "Don't trust this", no matter how good the lift looks. If the test stopped too early, ran for less than a week or has too few conversions, a winning result becomes "Keep running" until the test is done. A clean test with a small lift is worth more than a big lift on broken data.
In 7 of the last 10 analytics setups we audited, the tracking was broken in some way. In our experience, broken tracking leads to more bad decisions than bad maths. If we set up your testing and tracking, these checks happen before any result reaches your dashboard.
Why do winning A/B tests often shrink after launch?
Because the tests that just scrape through as winners are often the lucky ones. This is called the winner's curse.
Every result is a mix of the real effect and random noise. When a test only just passes, the noise has usually pushed the number up. Once you launch the change, the noise goes away, and the lift drops back towards its real size.
This happens most with small tests. Two habits help:
- Report the range, not just the headline lift.
- Base revenue forecasts on the low end of the range.
If a result only just passed, treat the lift as a best case, not a promise. The calculator warns you when a winning result only just passed.
How does this A/B test calculator work?
It uses a two-proportion z-test. It's the standard way to compare two conversion rates, described in the NIST Engineering Statistics Handbook. Here's the maths in plain terms:
- Conversion rate = conversions ÷ visitors, for each version
- Lift = (B's rate − A's rate) ÷ A's rate
- The test combines both versions' data to measure how much the rates would normally bounce around
- It then checks how far apart A and B are compared to that normal bounce, and turns that into a p-value
- The chance that B beats A starts with no opinion about either version, then updates on your data
Calculators make slightly different choices in these steps, which is why two tools can show slightly different numbers for the same data.
A/B test calculator FAQs
Should I use a one-sided or two-sided test?
Use two-sided by default. It tells you if B is better or worse than A. Use one-sided only if you decided before the test that you'd ship B only if it beat A. You can switch under Test details.
Is 90% confidence good enough for an A/B test?
It can be, for small changes that are easy to undo, like copy or images. For pricing, checkout or anything hard to reverse, use 95% or 99%. Pick your level before the test starts, not after you see the results.
What should I do if my A/B test isn't significant?
If you reached your planned number of visitors, the change probably makes less difference than you hoped. Keep A, or pick B for other reasons. The calculator also tells you the biggest lift the test rules out, so you know what size of effect you could have missed. If you haven't reached your planned visitors yet, keep the test running.
Can I test more than two versions at once?
Yes, but not in this calculator, which compares two versions. Testing A, B, C and D at once raises the chance of a false winner, unless you adjust for it. Each version needs its own full set of visitors, so an A/B/C/D test takes at least twice as long as an A/B test.
Can I use this calculator for revenue or average order value?
Not yet. This calculator compares conversion rates, which are yes-or-no outcomes: someone either converted or didn't. Revenue needs a different test, because a few big orders can skew the average. Use this calculator for sign-ups, purchases, clicks and similar actions.