A/B Test Your Emails Properly

Designs an email test with one variable, enough volume to resolve, and a decision rule set in advance — and tells you when your list is too small for testing to ever produce an answer. Use it before running a test you'll misread.

0 likes 0 dislikes
Sign in to rate this prompt

Prompt

    You are an experimentation-literate marketer. You are more concerned with whether a test can produce a real answer than with running one.

What I want to test: {{test_idea}}
What I think will happen and why: {{hypothesis}}
My list size and the segment being tested: {{list_size}}
Current performance on the metric in question: {{baseline}}
How often I send this type of email: {{send_frequency}}
What decision the result would inform: {{decision}}

Produce:

**Pick the right metric.** If {{test_idea}} is a subject line, the instinct is to measure open rate — don't. A large majority of opens are now logged automatically by mail clients that pre-load content, which inflates the number and makes it insensitive to what the subject line actually did. Measure click-to-open, clicks, or conversions instead. Say which metric this test should use and why.

**Sample size reality check.** Given {{list_size}} and {{baseline}}, roughly what size of difference could this test actually detect? Show the reasoning. Most email tests on lists under a few thousand can only detect very large differences, which means a small real improvement will register as no result — and a big-looking difference on a small sample is usually noise. If the test can't resolve, say so plainly and give me the alternatives: test on the highest-impact variable only, accumulate results across several sends, test something bolder, or make the change on judgment and monitor.

**One variable.** From {{test_idea}}: isolate it. Testing a different subject line *and* a different send time produces a result you can't attribute. If there are several things to test, sequence them and say in what order.

**Test something meaningful.** Small wording tweaks rarely produce detectable differences. Rank by expected effect size: the offer, the audience or segment, the format or structure, the call to action, the send time, and finally the wording. Most teams test the last one and wonder why nothing moves. Say where {{test_idea}} sits and whether it's worth a test at all.

**Design.** Split method, group sizes, whether to hold back a portion for the winner, and — if using a send-then-winner approach — the caution that declaring a winner on early opens rewards whoever gets opened fastest rather than whoever performs best.

**Decision rules, pre-committed.** What result means adopt, what means reject, and what means inconclusive. Write them now. Also the rule against checking early and stopping on a good-looking hour.

**Confounds.** Send day and time, seasonality, a concurrent campaign, and the fact that one variant may reach a systematically different audience if the split isn't random.

**What a win would actually be worth.** Against {{send_frequency}} and {{decision}}: if this improves the metric by a plausible amount, what does that mean over a year? Some tests aren't worth the effort even if they work, and knowing that in advance is worth more than the result.

**What to learn regardless.** The observation worth recording even if the test is inconclusive, so the next one is better designed.

Like this prompt?

Create an account to copy this prompt, create your own, and find the best prompts to scale your business.