About 73% of experimenters stop a test the moment it hits 90% confidence, and roughly 75% of measured effects are truly null. Sixteen prompts built against that.
Most of your ideas will not work, and you will conclude that they did
Two findings define this discipline. Neither is comfortable, and together they explain why so much conversion work produces confident reports and flat revenue.
The first: most ideas fail. Ron Kohavi, who ran experimentation at Microsoft, put the internal numbers plainly in his 2015 KDD keynote. Across experiments at Microsoft, one third of ideas were positive and statistically significant, one third were flat, and one third were negative and statistically significant. He adds two qualifications that matter: At Bing, the success rate is lower
— because a heavily optimised surface has less room left — and the low success rate has been documented many times across multiple companies.
Read that again: a third of the changes teams believed in made things worse. Not neutral. Worse. Shipping on conviction is not a neutral act.
The second: the way tests are actually stopped manufactures winners. In p-Hacking and False Discovery in A/B Testing, Ron Berman, Leonid Pekelis, Aisling Scott and Christophe Van den Bulte examined 2,101 commercial experiments on the Optimizely platform, using a regression discontinuity design to detect stopping behaviour.
About 73% of experimenters stop the experiment just when a positive effect reaches 90% confidence. Not at a pre-planned sample size — at the moment the number looks good.
And the base rate is brutal: approximately 75% of the effects are truly null. Improper optional stopping raises the false discovery rate from 33% to 40% among tests p-hacked at 90% confidence. They estimate the expected cost of a false discovery at a 1.95% loss in lift, because a fake winner stops you looking for a real one.
Put the two together. Most ideas do nothing, three quarters of measured effects are noise, and the standard industry practice of watching a dashboard until it turns green converts that noise into a case study.
That is what this pack is built against. Almost every prompt here exists to slow down one specific step where the discipline usually breaks.
Earn the right to test at all
Decide Whether You Have Enough Traffic to A/B Test belongs first, and it is the prompt most likely to tell you something you do not want to hear. It runs the sample-size and duration arithmetic and says honestly when your traffic cannot support a test — which, for most sites, is most of the time.
This is not a reason to despair; it is a reason to stop pretending. A site without the traffic to detect a 10% effect should be making large, evidence-led changes and measuring them over long periods, or doing qualitative research, rather than running underpowered tests that produce a random walk of confident conclusions.
Get evidence before you get ideas
Given a 75% null rate, the return on better hypotheses is enormous. The way to raise the hit rate is to stop generating ideas from best-practice lists and start generating them from what is actually going wrong.
Run a Conversion Research Sprint gathers from analytics, session recordings, surveys, support tickets and sales calls, then turns the pile into a ranked list of problems. It is the highest-leverage prompt on this page, because everything downstream inherits the quality of what goes in here.
Find the Objections Killing Your Conversions mines the same material for unspoken doubts and maps each to where it should be answered. Objections are the most reliable source of test ideas that work, because a real person already told you.
Run a Usability Test on Five People is qualitative and unreasonably effective — a complete moderated test with a script that avoids leading the participant. Five people will not give you statistical significance and will routinely show you a blocker you would never have hypothesised.
Do a Heuristic Conversion Walkthrough of Any Page reviews against relevance, clarity, value, friction, distraction and anxiety. It is a structured way to look, not evidence — use it to generate candidates, not conclusions.
Then be disciplined about what you test
Write a Testable Hypothesis From Evidence converts an idea into a stated cause, a predicted effect, a success metric and a kill condition. That last element is the entire defence against the Berman finding: a test with a pre-committed stopping rule cannot be stopped when the number looks good, because the rule was fixed before you looked.
Prioritize Your Conversion Test Backlog scores by impact, confidence in the evidence, and cost to build — and flags ideas scoring high on confidence with no evidence behind them, which is the polite name for a hunch.
Read A/B Test Results Without Fooling Yourself is the prompt this whole page argues for. Peeking, novelty effects, segment mining, and metrics that moved for unrelated reasons. If you adopt one thing here, adopt the rule that you read results once, at the pre-planned point.
A practical warning from the same Kohavi keynote, worth knowing before you trust any number: At Bing over 50% of traffic is bot generated.
If your test population includes bots, crawlers and internal traffic, you are measuring something other than customers.
Build a Conversion Test Log Your Team Learns From is how a two-thirds failure rate becomes an asset instead of an embarrassment. Given the base rates above, the losses are most of your data. A team that records only its wins re-runs its failures every eighteen months as staff turn over.
Fix the things that reliably lose money
Some changes are well-evidenced enough that testing them is a formality — and if you lack traffic to test, this is where to spend your effort.
Cart abandonment is the standing example. The Baymard Institute's average across 50 studies from 2006 to 2025 is 70.22%, a figure that has barely moved in two decades. Cut Friction Out of Your Checkout or Signup Flow works through the causes that research consistently identifies — surprise costs, forced account creation, long forms, unclear steps.
Fix Your Mobile Experience Where It Actually Loses Money targets speed on real devices, input and autofill friction, thumb reach and interruption. On performance, Kohavi's economics are striking: at Bing, an engineer improving server performance by 10 milliseconds — a thirtieth of the time it takes to blink — more than pays for his fully-loaded annual costs.
Speed is not a hygiene factor.
Write the Above-the-Fold Value Proposition drafts and stress-tests the headline, subhead and primary call to action, with a five-second test. Use Social Proof That Actually Persuades places the right kind of proof against each specific doubt, and strips the generic badges that occupy space without answering anything.
Design a Pricing Page That Helps People Choose treats the page as a decision aid rather than a display. See pricing.
For physical products, Write a Product Detail Page That Removes Doubt and Plan a Product Photo and Video Shot List both start from the questions that stop a purchase, so the imagery answers objections rather than decorating the page — more in ecommerce.
Where this fits
Conversion work sits directly downstream of paid advertising, and the two packs share a discipline: both are dominated by measurement error, and both reward people who are honest about what their data can and cannot support. Match Your Landing Page to the Ad's Promise is the seam between them.
For the underlying statistics — power, stopping rules, and whether a difference is real — see Check Whether a Difference Is Real or Just Noise and the analytics pack.
Where this stops
These prompts have no access to your analytics, your traffic, or your test results. They cannot run a significance calculation on data they cannot see, and they will not supply benchmark conversion rates — those vary so widely by industry, traffic source, price point and device that a generic figure is worse than no figure.
They also cannot tell you your tracking is broken, your test is contaminated by bots or internal traffic, or your randomisation is leaking across variants. Every method here will run cleanly on corrupted data and give you a tidy wrong answer.
And no amount of conversion optimisation fixes an offer people do not want. If the page is honest and the traffic is qualified and it still does not convert, the problem is upstream in the product, the price, or who you are attracting — see Find the Objections Killing Your Conversions first, and then consider that the answer may not be a page change at all.
Sources
- Ron Kohavi, Online Controlled Experiments: Lessons from Running A/B/n Tests for 12 years, KDD 2015 keynote, Microsoft — of ideas tested at Microsoft, one third positive and significant, one third flat, one third negative and significant;
At Bing, the success rate is lower
; over 50% of Bing traffic is bot generated; a 10ms server performance improvement more than pays for an engineer's fully-loaded annual cost - Ron Berman, Leonid Pekelis, Aisling Scott and Christophe Van den Bulte, p-Hacking and False Discovery in A/B Testing, December 2018 — 2,101 commercial experiments on Optimizely; about 73% of experimenters stop just as a positive effect reaches 90% confidence, approximately 75% of effects are truly null, optional stopping raises the false discovery rate from 33% to 40%, and the expected cost of a false discovery is a 1.95% loss in lift
- Baymard Institute, Cart Abandonment Rate Statistics — average documented online shopping cart abandonment rate of 70.22%, calculated across 50 different studies published between 2006 and 2025