In many teams the A/B test is presented as a referee that ends arguments: if we disagree, let us test it. Yet tests are not cheap. They spend traffic, time and attention. The real skill is not running a test but choosing which decisions deserve one.
Nielsen Norman Group points out that A/B testing tells you which option performs better but not why it does, and that understanding the reason requires qualitative research. An experiment therefore complements user research rather than replacing it.Nielsen Norman Group — A/B Testing 101
1. Not Every Decision Needs an Experiment
Some decisions already have a known answer. You do not test an accessibility defect, a broken form or an unreadable contrast ratio; you fix them. Putting those through an experiment only produces delay.
Other decisions are too small to measure. Testing a button colour on a page that receives a few hundred visitors a month takes months to reach a meaningful result and occupies the team's attention the whole time.
Decisions worth an experiment share three traits: the outcome is genuinely uncertain, there is enough traffic, and the result will steer an investment. If any one is missing, judgement, research or simply shipping is more efficient.
An experimentation culture built without that filter collapses under its own weight. As inconclusive tests pile up, the team starts believing that data does not work, when in fact the problem was the choice of question.
When defending experimentation, say the other part out loud too: not testing is not laziness. Re-measuring a question that established usability principles already answered steals time from the uncertainties that matter.
2. Match the Decision to a Method
The type of uncertainty determines the right method. If you do not know what users fail to understand, you need observation rather than an experiment; if you are torn between two good options, an experiment is the right tool.
The table below maps the situations teams tend to confuse onto methods. The point is not to belittle testing but to isolate the questions it can answer.
Write your traffic threshold into the table as well. The same question may be answered by observation on a page with two hundred daily visitors and by an experiment on one with twenty thousand.
| Situation | Type of uncertainty | Suitable method | Expected output |
|---|---|---|---|
| Users abandon the form | Reason unknown | Session observation and interviews | The friction point |
| Torn between two headlines | Preference unclear | A/B test | A measured difference |
| Broken flow or accessibility defect | No uncertainty | Direct fix | The defect removed |
| Idea for a new page type | Scope unclear | Prototype and usability test | Comprehensibility |
| Small change on a low-traffic page | Insufficient statistical power | Judgement and principles | A fast decision |
3. Write the Hypothesis and Metric First
A good hypothesis has three parts: for whom, which change, and what behavioural difference you expect. A test started without those three can be interpreted whichever way the numbers fall.
Choose a single primary metric. Teams that watch several at once build a story around whichever one moved, which turns the test into a confirmation ritual.
Put a guardrail metric next to the primary one. If a change lifts conversion while raising returns or support requests, the gain is not real.
Write the decision before launch as well: which result ships, which result rolls back, and which result counts as insufficient data. Without those three thresholds on paper, the outcome is always read in your favour.
- State the hypothesis in one sentence.
- Pick a single primary metric.
- Define the guardrail metric.
- Calculate sample size and duration up front.
- Record the decision thresholds in advance.
4. Hypothetical Scenario: The Winning Test That Lost Money
On a hypothetical services site the team tries a variant that cuts the quote form from seven fields to three. Two weeks later form submissions are clearly up and the variant is declared the winner.
A month later the sales team complains: most incoming requests come from people whose budget or scope does not fit, and the rate of progressing to a first meeting has dropped. The test won and the business lost.
The mistake sits in the metric, not the test design. Had the primary metric been qualified meetings rather than form submissions, with time-to-sale per request as a guardrail, the result would have read differently from the start. This is a hypothetical example, not a client result.
5. Statistical Discipline
Reliability usually suffers from impatience rather than from statistics. Checking results daily and stopping the moment significance appears inflates the false positive rate considerably.
That is precisely the warning Evan Miller is known for: repeatedly checking results and stopping as soon as significance shows pushes the true error rate far above the one you announced, and the remedy is to fix the sample size in advance and hold to it. An experiment calendar therefore follows statistical power, not the marketing calendar.Evan Miller — How Not To Run An A/B Test
Set the duration against your business cycle too. On a site where weekday and weekend behaviour differ, a three-day test measures the calendar rather than the change.
When the result is inconclusive, do not treat it as failure. Deciding not to ship a change with no measurable effect is also a decision, and it moves the team on to a bigger question.
Calculate the sample in advance
Before the test, write down the current conversion rate, the smallest difference worth detecting and the acceptable error margin. Those three values determine the number of visitors and the duration you need.
If the calculated duration exceeds three months, do not run the test. That is a sign the question needs judgement or a larger change instead.
Avoid the segment trap
When the overall result is flat, slicing into segments to find a winning sliver is the most common way to mistake noise for discovery. Segment analysis is only trustworthy when defined beforehand.
If an unexpected segment difference looks interesting, record it as a hypothesis rather than a result and test it separately.
6. Building the Experimentation System
Establish the process before choosing a tool. A testing platform bought without a hypothesis backlog, prioritisation criteria, decision thresholds and a results archive only produces inconclusive answers faster.
Collect hypotheses in one list and rank them by expected impact, implementation cost and statistical power. Once the ranking criteria are written down, the debate shifts from opinion to method.
Record every result together with what you learned, before labelling it a win or a loss. Six months later your most valuable asset is not the winning variants but the archive of ideas already tried and eliminated.
Keep that archive searchable. Each entry should carry the hypothesis, audience, duration, sample, primary metric, result and decision. With standard fields, a new team member will not retry the same idea six months later.
- Hypothesis backlog collected in one list.
- Prioritisation criteria written down.
- Primary and guardrail metrics chosen.
- Sample size and duration calculated up front.
- Decision thresholds recorded.
- Test limited to a single variable.
- Result and learning written to the archive.
7. Limits and Failure Modes
An A/B test does not answer why. It tells you which variant won; only observation and interviews reveal what users failed to understand or why they hesitated.
The gains accumulated by small changes are also bounded. A programme that advances only through colour and copy trials eventually finds itself chasing differences too small to measure.
Novelty effects can inflate a result temporarily. Existing users' first reaction to a new layout may not represent long-term behaviour, so keep watching the metric after critical changes.
Experiments have ethical and legal limits too. Trials involving pricing, contract terms or the scope of personal data collection can be unfair to users or non-compliant; those areas need a legal assessment first.
Finally, an experimentation system is a cultural commitment. In a team that accepts results only when they confirm expectations, testing becomes an instrument of persuasion rather than decision.
Conclusion
An experimentation system is built by separating the right questions, not by running more tests. Match decisions to methods, write hypotheses and thresholds in advance, calculate the sample and archive what you learn.
Frequently Asked Questions
Sources
- Nielsen Norman Group — A/B Testing 101
Which question testing answers and which it does not
- Evan Miller — How Not To Run An A/B Test
How early stopping affects the error rate
Turn your experimentation into a system that decides
Let us set up your hypothesis backlog, metrics and decision thresholds together.
Request an experimentation setup


