Conversion Research

What an A/B test measures, and how to read the results

Statistical significance in plain terms, for founders who see test reports.

Somebody sends you a report. The new page version won by 18%. The result is described as significant. There is a green box.

Before acting on it, it is worth knowing what that claim asserts, because in B2B the answer is frequently “less than you think”, and occasionally “nothing at all”.

Statistical significance says that results this extreme would be uncommon if the change did nothing. It does not say the result is real, and it says nothing about the size of the effect.

In this article

What an A/B test is doing

An A/B test splits visitors randomly between two versions of a page and compares an outcome, usually a conversion rate.

The random split is the entire value of the method. Because assignment is random, the two groups are on average alike in every respect except the change you made. Any reliable difference in outcome can then be attributed to the change rather than to who happened to see which version.

That is a strong tool, and it is why large technology companies run enormous numbers of these tests. The authors of Trustworthy Online Controlled Experiments, who led experimentation at Microsoft, Google and LinkedIn, describe organisations running more than twenty thousand controlled experiments a year.

The important word in that sentence is twenty thousand. It is a clue about what makes the method work, and about why it often does not work on a B2B site.

What “significant” means

Statistical significance is a statement about a specific hypothetical.

It says that if the change had no real effect, results at least this extreme would occur only rarely by chance. The conventional threshold is 5%, usually written as p < 0.05.

Three things follow, and each is routinely misread.

It is not the probability that the result is real. It is the probability of seeing data this extreme in a world where the change did nothing. Those are different statements.

It says nothing about the size of the effect. With enough traffic, a trivial improvement becomes significant. Significance tells you the direction is probably not noise. It does not tell you the difference is worth having.

It carries a built-in error rate. At a 5% threshold, one in twenty tests of a change that does nothing will still report a win. Run twenty pointless tests and expect roughly one false victory, reported confidently, with a green box.

The mistakes that produce wrong answers

Kohavi and colleagues catalogue the failure modes. Four of them account for most bad B2B test reports.

Stopping when it looks good. This is called peeking. If you check the results daily and stop the moment significance appears, you have not run a test with a 5% error rate. You have run a procedure that finds a “winner” for changes that do nothing far more often than that. The fix is to decide the duration before starting and read the result at the end.

Not enough traffic. Detecting a small difference requires a large sample. A page with a few hundred visits a month cannot reliably detect a 10% change in conversion. It can produce a number, and the number will move around, and none of that movement means anything.

Testing many things and reporting the winner. If you test six variations, the chance that at least one clears the threshold by luck is much higher than 5%. Reporting only the winner conceals the other five attempts.

Running for less than full business cycles. Traffic behaves differently on weekdays and weekends, and B2B traffic behaves very differently. A test run Tuesday to Friday has sampled one kind of visitor.

The B2B traffic problem

Most B2B websites do not have enough traffic to run valid A/B tests on their main conversion.

This is arithmetic, not pessimism. To detect a meaningful change in enquiry rate you need enough conversions in each group for the comparison to mean anything. A site receiving forty enquiries a month, split across two versions, produces twenty per group. At those numbers the difference between the groups is dominated by chance.

The honest conclusion is that for a large share of B2B companies, A/B testing the enquiry form is not a viable method. Running it anyway produces confident-looking reports that are indistinguishable from coin flips, and decisions get made on them.

Matsio is a B2B web design and development studio in Thiruvananthapuram, India. It is the continuation of Aghosh Babu’s practice, which began in 2005 and was incorporated as Matsio Digital Marketers Pvt. Ltd. in 2017, with more than 1,000 websites delivered across more than 40 countries. On sites without the traffic to support valid experiments, the studio’s position is to say so rather than to produce a test report that cannot support its own conclusion. 85% of clients return for further work, measured across every engagement since 2005, and being told when a method does not apply is part of why.

What to use when testing is not available

The absence of statistical testing does not mean flying blind. It means using methods appropriate to the sample size you have.

  • Usability testing with five to eight people. This finds problems, not percentages. If four of six people cannot find your pricing, you do not need a significance calculation to act.
  • Session recordings. Watching twenty real sessions on your enquiry form will usually reveal more actionable problems than a test you cannot power.
  • Sequential comparison with honest caveats. Change the page, watch the following quarter, and record that this is not a controlled comparison. Seasonality and campaigns are uncontrolled. It is weak evidence, and weak evidence labelled as such is more useful than strong-looking evidence that is wrong.
  • Testing higher up the funnel. You may not have enough enquiries to test, but you may have enough clicks or scroll events. Testing an intermediate step can be validly powered when the final step is not.
  • Applying established research. Where the evidence base is large and consistent, such as form field research or page speed research, adopting the finding is more reliable than an underpowered test of it on your own site.

What a valid test looks like when you can run one

Some B2B sites do have the traffic, usually on high-volume entry points rather than on the final conversion. When you can run a test properly, the discipline is short and non-negotiable.

Decide what you are testing and why, in one sentence, before you start. A test without a hypothesis is a search for anything that moves, and something always moves.

Calculate the sample size in advance. This requires your current conversion rate and the smallest improvement you would act on. The calculation tells you how long the test must run. If the answer is fourteen months, you have learned something valuable before spending anything.

Fix the duration and do not look at significance until it ends. Monitoring for technical breakage is fine. Monitoring for a winner is peeking.

Run for whole weeks, so both halves see the same mix of weekday and weekend traffic.

Report the confidence interval alongside the point estimate, and report every variation you ran.

Then, if it wins, consider running it again. Replication is the least popular and most informative step in the whole method.

Who is asking for the test, and why

There is an organisational dimension to this that is usually the real issue.

Requests for A/B tests often arrive when a team disagrees about a change and wants an objective tiebreaker. That is a reasonable instinct and a poor use of the method, because the test will produce a number regardless of whether it can support one, and the number will settle an argument it was never powered to settle.

The better response to a disagreement is usually to establish which of the two positions has research behind it, and to be explicit when neither does. A decision made on stated reasoning that turns out wrong is recoverable. A decision made on a misread test is defended, because it has a number attached.

How to read the next test report you are sent

Five questions, in order.

How many conversions were in each group? Not visitors, conversions. Under a few hundred per group, treat the result as provisional whatever the report says.

How long did it run, and was the duration decided in advance? If it was stopped when it looked good, the significance figure is not valid.

How many variations were tested, and how many were reported? Ask what else was tried.

What is the confidence interval, not just the point estimate? “18% better” means little. “Somewhere between 2% worse and 35% better” means the test did not resolve the question.

Does the effect size matter commercially? A significant 0.4% improvement on a page with low traffic may not be worth the deployment risk.

The short version

Statistical significance means results this extreme would be uncommon if the change did nothing. It does not mean the result is real, does not describe the size of the effect, and is invalidated by stopping early or testing many things without reporting them.

Most B2B sites lack the traffic to test their main conversion validly. Knowing that is more valuable than a test report that cannot support its conclusion, because the report will be believed and acted on regardless.

A small thing and a big thing

One small thing to fix on your website today, and one big thing to learn that gets you more leads.

logo
© 2026 Matsio Digital Marketers Pvt. Ltd.