Across 663 Facebook experiments, true lower-funnel lift was 5% while observational attribution reported 24-64%. Build from what you can afford to pay, not from the dashboard.
The number in your ads dashboard is not the return on your ad spend
This pack was written last, and deliberately, because the honest version of paid advertising rests on a body of evidence that most of the industry works hard not to look at.
Start with the largest study of the question. In Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement (Marketing Science, 2023), Brett Gordon, Robert Moakler and Florian Zettelmeyer analysed 663 large-scale randomised experiments at Facebook, with access to over 5,000 user-level features — far richer data than any advertiser or measurement partner can obtain.
They then asked whether the best available observational methods could recover the true experimental result. The answer, in their numbers:
| Funnel stage | True lift (RCT) | DML estimate | Matching estimate |
|---|---|---|---|
| Upper | 29% | 83% | 173% |
| Middle | 18% | 58% | 176% |
| Lower | 5% | 24% | 64% |
At the lower funnel — purchases, the thing you are actually buying — the true median lift was 5%, and observational attribution reported somewhere between 24% and 64%. Overstatement of five to thirteen times, using better data than you have.
Their conclusion is worth quoting exactly: despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign's causal effect.
What that means concretely
The mechanism is selection. Ad platforms optimise delivery toward people most likely to convert, so the people who saw your ad were already the most likely buyers. Attribution then credits the ad with purchases that were going to happen anyway.
The cleanest demonstration is Thomas Blake, Chris Nosko and Steven Tadelis's eBay field experiments, published in Econometrica in 2015. Turning off paid search revealed that brand-keyword ads have no measurable short-term benefits
— people searching for eBay found eBay. Returns from paid search overall were a fraction of conventional non-experimental estimates.
Worse, the spend concentrated on frequent users whose behaviour the ads did not change, producing negative average returns for that segment while the dashboard reported success.
So the money went to the customers who needed no persuading, and the reporting confirmed the decision. That is not a measurement bug at the margin. It is the central failure mode of the channel.
And measuring it properly is genuinely hard
Before assuming a holdout test will settle everything: Randall Lewis and Justin Rao's The Unfavorable Economics of Measuring the Returns to Advertising (Quarterly Journal of Economics, 2015) ran 25 large field experiments with major US retailers and brokerages, totalling $2.8 million in ad spend.
The median confidence interval on ROI was over 100 percentage points wide. Individual-level sales are extremely volatile — a coefficient of variation of 10 is common — so an informative experiment can easily require more than 10 million person-weeks.
The implication is not give up.
It is that small advertisers cannot statistically detect moderate effects, and any confident ROAS figure to two decimal places is describing precision that does not exist. Design your measurement around decisions that survive that uncertainty.
So build from the floor, not the dashboard
If the reported return is unreliable, the number you must own is what you can afford to pay.
Set Your Target CPA and ROAS From Unit Economics is the foundation of this pack for that reason. Margin, repeat purchase and payback period give you an affordable acquisition cost that is true regardless of what any platform reports. Pick a target ROAS by benchmark and you have anchored your business to a number someone made up.
Build a Paid Media Plan and Allocate Budget then allocates against it, with a reserve held for testing and stated conditions for shifting money — decided in advance, when you are calm.
Measure Paid Performance When Attribution Is Broken is the direct response to the research above: holdout tests, blended metrics and simple incrementality, rather than platform-reported conversions. Geo holdouts and full-channel pauses are crude, and crude and causal beats precise and biased.
Given Lewis and Rao, size these honestly — Decide Whether You Have Enough Traffic to A/B Test from the CRO pack runs the same arithmetic and will sometimes tell you the test you want is not available to you.
Build the account so it can be read
Structure a Paid Campaign Account Properly matters more now than when you controlled targeting manually, because the platform's automation performs according to how you have grouped and signalled — and because an unreadable account cannot be diagnosed later.
Build Audience Targeting and Exclusion Lists is where the eBay finding becomes operational. Exclusions are the lever. Existing customers, recent purchasers and people already arriving through organic channels are exactly the population that generates impressive attributed returns and no incremental revenue.
Build a Retargeting Strategy That Isn't Creepy is the same argument at its sharpest. Retargeting reports the best ROAS in most accounts, and structurally must: you are advertising to people who already chose you. Segment by behaviour, cap frequency, bound the window, and treat the reported figure with suspicion proportional to how good it looks.
Fix a Product Feed That Keeps Getting Disapproved is unglamorous and, for shopping campaigns, often the single largest recoverable loss. See ecommerce.
Creative is the variable you still control
As targeting has moved inside the platforms, creative has become the main lever an advertiser actually operates.
Brief Paid Social Creative That Stops the Scroll specifies the job of the first three seconds, because that is where the decision happens. Write Search Ad Copy Variations Worth Testing produces genuinely different variants rather than reworded twins — testing synonyms consumes budget and teaches nothing.
Build a Creative Testing System is the one that compounds: a repeatable process with stated rules for what counts as a winner and when to retire a tiring asset. Note the statistical caution from the CRO pack applies here in full — Read A/B Test Results Without Fooling Yourself, because creative tests are stopped early more often than any other kind.
Match Your Landing Page to the Ad's Promise covers the most common failure in the whole funnel. Perfect targeting into a page that answers a different question converts nobody, and the account will look like a media problem.
Diagnose, and stop the leaks
Cut Wasted Ad Spend finds irrelevant search terms, junk placements, overlapping audiences and campaigns nobody has looked at since launch. It is usually the fastest positive return available in an established account.
Diagnose a Campaign That Stopped Performing works the causes in order — creative fatigue, auction shifts, tracking breakage, audience exhaustion, someone's change — instead of jumping to the most emotionally available explanation, which is usually the algorithm.
Audit a Paid Account You Inherited starts with tracking, because everything downstream is meaningless if the measurement is wrong.
Test a New Channel Before You Commit Budget sets a learning goal, enough budget for a fair trial, and a decision rule fixed in advance — which is the only defence against a test that runs until someone likes the numbers.
Where this stops
Nothing here reports on your account. These prompts have no access to your platforms, your spend, your conversions or your margins, and they will not produce a benchmark CPM, CPC, CTR or industry-average ROAS — those figures vary enormously by geography, vertical, season and auction, and any specific number generated for you would be invention. Where you need a benchmark, get it from your own historical data.
Ad platforms' policies, prohibited categories and disclosure requirements change constantly and differ by jurisdiction and vertical — finance, health, housing, employment and credit especially, where targeting is legally restricted. Check the current policy rather than a prompt's recollection of it, and take advice where advertising law applies to your category.
Sources
- Brett R. Gordon, Robert Moakler and Florian Zettelmeyer, Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement, Marketing Science 42(4), 2023, 768–793 — 663 large-scale Facebook experiments with 5,000+ user-level features; median RCT lifts of 29%/18%/5% by funnel stage against DML estimates of 83%/58%/24% and matching estimates of 173%/176%/64%; the authors conclude they are
unable to reliably estimate an ad campaign's causal effect
- Thomas Blake, Chris Nosko and Steven Tadelis, Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment, Econometrica 83(1), 2015, 155–174 — eBay field experiments; brand-keyword ads have no measurable short-term benefits, returns are a fraction of non-experimental estimates, and spend concentrated on frequent users produced negative average returns
- Randall A. Lewis and Justin M. Rao, The Unfavorable Economics of Measuring the Returns to Advertising, Quarterly Journal of Economics 130(4), 2015, 1941–1973 — 25 large field experiments totalling $2.8m in spend; median ROI confidence interval over 100 percentage points wide, and informative experiments can require more than 10 million person-weeks