To A/B test cold email, pick one change, measure positive replies per lead, and work out the sample size before you send. At a 2% baseline, detecting a lift to 3% takes about 3,825 leads per version, nearly 7,700 in total. Most cold programs do not have that volume per test, so test big changes, not small ones, and accept “no difference” as an answer.
This article gives the reasoning, a worked calculation you can check, and the mistakes that produce a false winner. The statistics are the standard two-proportion test. The numbers below are computed from the formulas shown, and the example rates are assumptions, not benchmarks.
What to measure
Use positive replies. A positive reply is a response that shows interest: a yes, a request for information, a referral to the right person. Count it per lead, within a fixed window.
| Metric | Use in a test? | Why |
|---|---|---|
| Positive reply rate | Yes, the decision metric | A person acted on the offer |
| Any reply rate | As a second view | Includes “no” and “stop”, and catches messages that provoke |
| Meeting booked | If you have the volume | The result you want, but too rare to test with small samples |
| Opens | No | Image loading is automatic for many recipients, see open rates |
| Bounce, complaint, unsubscribe | As guardrails | A winner that doubles complaints is not a winner |
Define the rules before the test starts:
- The unit is the lead. One lead gets one variant for the entire sequence, so one person never sees both versions.
- The window is fixed. For example, replies within seven days of the last step. Compare variants over the same window.
- Classify replies the same way. Decide what counts as positive in writing, and ideally label replies without looking at the variant.
- Exclude out-of-office and bounces from both numerator and denominator, or at least treat them the same in both arms.
Test one thing
If you change the subject, the first line and the call to action at once, you learn that version B beat version A and nothing about why. Pick one element: the subject, the opening line, the offer, the call to action, or the length. Large differences are easier to detect than small ones, which matters because of the sample sizes in the next section.
The sample size, step by step
For two proportions, p1 (the baseline) and p2 (the rate you want to detect), with a two-sided test at significance level alpha and power 1 - beta, the standard normal-approximation formula for leads per variant is:
n = ( z_a * sqrt(2 * pbar * (1 - pbar)) + z_b * sqrt(p1*(1-p1) + p2*(1-p2)) )^2 / (p1 - p2)^2
pbar = (p1 + p2) / 2
z_a = 1.96 for alpha = 0.05, two-sided
z_b = 0.8416 for power = 80%
Worked example. Baseline p1 = 2%, target p2 = 3%:
pbar = 0.025
2 * pbar * (1 - pbar) = 0.04875, sqrt = 0.22079, times 1.96 = 0.43276
p1*(1-p1) + p2*(1-p2) = 0.0196 + 0.0291 = 0.0487, sqrt = 0.22068, times 0.8416 = 0.18572
sum = 0.61848, squared = 0.38252
(p1 - p2)^2 = 0.0001
n = 0.38252 / 0.0001 = 3,825 leads per variant
That is 7,650 leads for two variants. We checked the number with a simulation: with true rates of 2% and 3% and 3,825 leads per arm, 80.1% of 20,000 simulated tests were significant at the 5% level, as the formula says.
More cases from the same formula, same alpha and power:
| Baseline | Target | Leads per variant | Both variants |
|---|---|---|---|
| 1% | 2% | 2,319 | 4,638 |
| 2% | 4% | 1,141 | 2,282 |
| 2% | 3% | 3,825 | 7,650 |
| 3% | 4.5% | 2,517 | 5,034 |
| 3% | 4% | 5,301 | 10,602 |
| 5% | 6% | 8,158 | 16,316 |
Notice the pattern. Doubling the rate needs a thousand or two per arm. Adding one point on a 5% base needs eight thousand. The smaller the lift you want to see, the more leads you need, and it grows with the square of the difference.
What a small test looks like
Suppose you send each of two versions to 1,000 leads and get 20 positive replies on A and 30 on B. B looks 50% better. The two-proportion z-test (the large-sample method described in the NIST handbook) gives:
pA = 0.020, pB = 0.030, pooled p = 50 / 2,000 = 0.025
z = (0.030 - 0.020) / sqrt(0.025 * 0.975 * (1/1000 + 1/1000)) = 1.43
two-sided p-value = 0.15
A p-value of 0.15 means a difference at least this large would appear about 15% of the time even if A and B were identical. It is not significant. Another view is the 95% Wilson interval for each rate: about 1.3% to 3.1% for A and 2.1% to 4.3% for B. They overlap heavily.
The test could not have worked anyway. With 1,000 leads per arm and true rates of 2% and 3%, our simulation detected the difference in only 30% of runs. That is the power of the test, and it is why the answer to a small test is usually “we cannot tell”, not “B won”.
How long a test takes
Volume decides how many tests you can run. With the assumed program from the cost model, 300 emails a day at 2.5 emails per lead is 120 new leads a day. A 2% to 3% test needing 7,650 leads takes about 64 sending days, roughly three months, if the whole program is dedicated to it. A test of 2% against 4% needs 2,282 leads, about 19 days.
So a small program should:
- Test large changes, such as a different offer or a different target role, and not a word in the subject line.
- Run few tests, and run each to its planned size.
- Accept that many tests end without a winner. That is information too: it says the change does not matter much.
Mistakes that produce a false winner
- Peeking. Checking the result every few hundred sends and stopping at the first p-value under 0.05. We simulated five looks (every 1,000 leads per arm up to 5,000) with both versions at the same 2% rate: 13% of runs declared a winner, not the 5% you set. Fix the sample size and look once.
- Unequal conditions. If variant A went out Monday from newer mailboxes and variant B Wednesday from older ones, you measured the mailboxes and days. Split each variant evenly across mailboxes, domains and days, and assign leads at random.
- Different audiences. Assigning by list order, region or company size puts different people in each arm.
- Counting sends, not leads. A lead who got three messages is one observation, not three.
- Changing the test midway. Editing a version while the test runs starts a new test.
- Many variants and many comparisons. With four variants there are six pairs. The more pairs you check, the more likely one looks significant by chance, so require a stricter threshold, or plan the extra sample.
- Reading a winner too broadly. A result for one audience and offer says little about another.
- Ignoring the guardrails. A version that wins on positive replies and doubles unsubscribes or complaints costs you more elsewhere. Providers such as Google and Yahoo watch complaints against a 0.3% limit.
A procedure you can follow
- Write the hypothesis: “Version B, with the offer in the first line, raises positive replies.”
- Choose the baseline from your own history, and the smallest lift worth acting on.
- Compute n per variant with the formula, and check you can reach it in a reasonable time.
- Assign leads at random and evenly across mailboxes and days.
- Run to the planned sample plus the reply window. Do not look at the winner before.
- Compare positive reply rates with the two-proportion test, check the guardrail metrics, and write down the decision and the reason.
- Roll out the winner, and next test something else.
How Dooxout handles this
Dooxout lets you put two to four variants on a campaign step (subject, body or both), with a weight for each variant. Assignment is deterministic per lead, so recalculating never moves a lead to another variant, and results are read on positive replies first, then any reply, not on opens. A winner can be picked manually, by a rule that needs a minimum number of sends per variant and a significance threshold, or proposed by an agent for a person to approve. The platform blocks ending a test before the minimum sample and records the decision with a reason.
The built-in SDR agent can draft A/B variants of the first message and suggest tests, and it works within the autonomy level a person sets. As elsewhere, the platform warns and recommends, and you decide. See AI outbound sales, email deliverability and integrations.
Frequently asked questions
What metric should I use to A/B test cold email?
Positive replies per delivered message, counted per lead. It is a person acting on the offer. Opens are inflated by automatic image loading, Apple states that senders will see emails as opened regardless of whether the user read them, so a test decided on opens can pick the wrong winner. See why open rates lie.
How many emails do I need per variant?
It depends on your baseline rate and on the smallest difference you care about. For a baseline of 2% positive replies and a target of 3%, a two-sided test at 5% significance with 80% power needs about 3,825 leads per variant. For 2% against 4% it is about 1,141. Compute it before you launch, with your own numbers.
Can I stop the test when one version is ahead?
Not safely. Checking repeatedly and stopping at the first significant result raises the false positive rate far above the 5% you chose. In our simulation of five looks at equal rates, about 13% of tests declared a winner that did not exist. Fix the sample size in advance and read the result once.
Can I test more than two versions at once?
Yes, but each extra variant needs its own sample, and comparing several pairs raises the chance of a false winner. Test one change at a time, and keep to two or three variants unless you have the volume.
Sources
The external facts in this article were checked against these pages on . Provider limits and rules change, so check the current page before you rely on a number.