Skip to content
AI GUIDERPROAIGuiderPRO — Smarter Search. Better Growth.
Conversion

Six ways conversion tests produce confident wrong answers

A/B testing feels rigorous, which is exactly what makes a badly run test dangerous. The six failures that turn noise into a decision.

AIGuiderPRO9 min read

A/B testing carries the aesthetics of science — a control, a variant, a percentage. That is precisely why a badly run test is more dangerous than no test: it produces a confident number that survives challenge in a meeting.

Six failures account for most of it.

1. Stopping when it looks good

A variant goes ahead on day three, somebody calls it, the change ships.

Early results swing wildly on small samples. Watching until a variant is ahead and stopping there is not measurement — it is waiting for noise to flatter you, and it will do so eventually in any test.

Declare the sample size and end date before starting, write them down, and stop only when both are met.

If you would not have stopped the test had the control been ahead, you are not running a test. You are looking for permission.

2. Testing below the traffic threshold

Roughly 5,000 relevant sessions a month is the practical floor for a meaningful test on a primary conversion.

Below that, tests either never reach significance or reach it spuriously. Either way you are paying to read noise, and the confident output is worse than no output.

With low traffic, fix the obviously broken things and invest in demand. That is the honest answer, and it is why we turn down CRO engagements below that threshold.

3. Measuring the wrong outcome

A variant that lifts form submissions by 30% while lowering qualified conversations is a loss reported as a win.

Removing qualifying fields, softening the call to action or overpromising in the headline all increase volume and decrease quality. If the measured outcome is the form fill, all three look like successes.

Measure to qualified conversation. It requires CRM connection and patience, and it is the only way the number means anything.

4. Testing things that should just be fixed

A broken form does not need an experiment. Nor does a page taking eight seconds to load, or pricing nobody can find.

Testing obvious breakage burns weeks establishing something everyone already knew, and it consumes the traffic budget a genuinely debatable test would have needed.

Test the changes where reasonable people disagree. Ship the rest.

5. Running overlapping tests

Two experiments on the same audience at the same time contaminate each other. When both show movement, neither result is attributable.

Sequence them, or split the audience properly with a tool built for it. Running them concurrently because time is short produces two unusable results rather than one usable one.

6. Ignoring the calendar

Buying behaviour differs by weekday. A test running Tuesday to Friday measures a different population from one covering a full week.

Cover at least one complete weekly cycle. If your business has monthly or seasonal patterns, cover those too, and never compare a test period against a structurally different one.

What a well-run test looks like

  • A written hypothesis naming what you expect to change and why.
  • A declared sample size and end date, fixed before launch.
  • One change tested at a time on one audience.
  • Success measured at qualified conversation, not form fill.
  • A full weekly cycle minimum.
  • The result recorded whether it won or lost.

Frequently asked questions

  • Until it reaches a pre-declared sample size and covers at least one full weekly cycle. Stopping when a variant looks ahead is the most common way tests produce confident wrong answers.

Structured data on this page

  • Article
  • FAQPage
  • BreadcrumbList

These schema types are implemented on this page. If you are applying this guidance to your own site, they are the ones worth deploying first.

Find out what AI says about you right now.

We run your real buyer prompts through four assistants and send you the transcript, with your position and your competitors’. No charge, no call required to receive it.

Typical turnaround: 3 working days.