Why attribution is not enough
Every ad platform hands you an attribution report. You spend money, someone buys, the platform takes credit. Run two platforms, and both claim the same sale. Run three, and you are paying for a fiction.
Attribution tells you where the platform saw a user before a conversion. It does not tell you whether the ad caused the conversion. Those are different questions, and conflating them has funded years of wasted spend.
Incrementality testing asks the causal question directly: if you removed this campaign, would the result have happened anyway? The answer comes from a controlled experiment, not from window-matching or weighted probability scores.
This matters more as budgets tighten. A board does not want to hear that your platform dashboard showed a 5x ROAS. They want to know whether removing that channel would change revenue. Incrementality is the method that produces an honest answer. MMM gives you the portfolio view; individual tests give you the isolated signal. They are complementary, not competing. For the relationship between them, see our post on why MMM has returned as a board-level measurement tool.
Three methods dominate. Each has a different set of trade-offs in accuracy, cost, and operational complexity.
The three methods: geo holdouts, conversion lift, and ghost bids
These methods differ in what they suppress, what they measure, and what assumptions they require you to accept.
Geo holdouts suppress advertising in a matched geography while running normally in a test geography. You measure the outcome difference between the two over the test window. The design is simple enough that you control it entirely without platform cooperation. Matched-pair design is the core requirement: you must select geographies that tracked each other closely on the relevant business metric before the test window began.
The holdout window varies by channel. Short-cycle campaigns (paid search) need at least four weeks for the effect to settle. Channels with longer adstock cycles (programmatic, CTV) typically need six to eight weeks before you can trust the delta. The existing leapbuzz practice on geo holdout calibration, documented in our post on cookieless MMM and Bayesian incrementality in APAC, covers channel-specific window guidance in detail.
Conversion lift studies use platform-side randomisation to split audiences into test and holdout cells. Users in the holdout cell are withheld from the campaign; their conversion behavior is compared to the test cell over the study period. Major social and programmatic platforms offer this natively, typically as a feature available to accounts above a volume threshold. The advantage is user-level precision across a single channel. The limitation is channel isolation: the holdout users may be reached by the same campaign on a different platform, contaminating the control.
Ghost bids are the cookieless answer to conversion lift. Instead of withholding users from the auction, the system submits phantom bids for users who would have been targeted, then tracks their downstream behavior without showing them the ad. The conversion delta between ghost-bid users and served users estimates the incremental effect. The method requires technical implementation at the platform or DSP level and is not universally available. It has grown in relevance as identity signals have narrowed.
Which method to start with
Geo holdouts require no platform relationship and can be run on any channel, including offline. Start there. Graduate to conversion lift when you need user-level precision on a single social channel and the platform relationship supports it. Use ghost bids when cookie-based user matching is unreliable in your market.
Designing a first test
Most incrementality tests fail at the design stage, not the measurement stage. Four decisions determine the validity of everything that follows.
Select the outcome metric before the test starts. Pick one. Revenue, new customers, store visits, app installs. Pre-registration matters because outcome switching after the test is a form of data manipulation that looks like analysis. If you are unsure which metric the business cares about, resolve that before designing the test.
Match geographies on the business metric, not population. Two cities of equal size are not a valid matched pair if one has twice the organic purchase rate. Match on the conversion rate or revenue-per-population of your specific outcome over a historical baseline of at least twelve weeks. Seasonal matching matters too: avoid testing windows that span major shopping events unless both geographies experience them identically.
Size the test for statistical power before running it. The calculator below gives you a quick power estimate. The practical implication: small budgets and low conversion rates require very large test populations or very long test windows to detect moderate lifts at acceptable confidence. If your test is underpowered, you will either declare a false negative (the campaign worked but the test was too small to show it) or make an irreversible budget decision on noise. A minimum detectable effect of 10 percent lift on a 2 percent baseline conversion rate requires far more traffic than most teams expect.
Lock down the suppression perimeter. For geo holdouts, map which campaigns run in the holdout geography before the test window. Any spend leaking into the holdout breaks the design. For conversion lift, confirm whether the holdout cell receives the same campaign on other platforms in your media mix. A user in your Meta holdout who is being served the same message on YouTube is not a clean control.
These four decisions made in advance will do more for result quality than any statistical sophistication applied after the fact.
Uplift significance calculator
Use this tool to estimate whether an observed lift is statistically meaningful or within noise range. Enter the visitor counts and conversions from your test and control groups. The calculator runs a two-proportion z-test and gives you a plain-language read on the result.
Incrementality uplift calculator
Planning model. Enter test and control group data to estimate lift significance.
Planning model only. Two-proportion z-test assumes independent groups, simple random assignment, and binary outcomes. It does not account for clustering in geo holdouts, covariate imbalance, or sequential testing inflation. Use it to size tests and sense-check results, not as a final statistical authority. For geo holdouts in particular, consult a statistician before acting on results near the significance boundary.
Reading results without fooling yourself
A positive lift number is not a pass. Three failure modes regularly inflate incrementality results in ways that look clean until you go looking for them.
Novelty effects. Users respond differently to a new campaign in its first week or two. Response rates are inflated by curiosity or freshness, then settle. If your test window is four weeks or shorter, the average lift may be driven by the opening burst rather than steady-state performance. A longer test window or a pre-analysis period that excludes the first week of exposure helps here. The effect is more pronounced for brand new creative than for refreshed campaigns.
Contamination. In geo holdouts, contamination enters when users cross your geographic boundary. A campaign suppressed in city B reaches users who drive, commute, or shop in city A. A user in your Meta holdout cell who sees the same campaign on a different platform is contaminated. Neither is a design failure you can always prevent, but you need to estimate its likely magnitude before reporting the result. High commuter corridors between test and holdout geographies are a known risk; select boundaries that align with commercial rather than residential geography where possible.
Selection bias in conversion lift cells. Platform-side holdout randomisation is not always truly random at the user level. Platforms assign holdout cells based on their own optimisation logic, which can systematically put lower-intent or harder-to-reach users into the control cell. If you see an unusually high conversion rate in your holdout group, that is a signal of potential cell imbalance, not evidence that your campaign hurt performance. Ask the platform for cell composition diagnostics where available.
One other read that trips teams up: a statistically significant lift at a small absolute size. A campaign that generates a 15 percent relative lift on a 0.3 percent baseline conversion rate has moved from 0.3 to 0.345 percent. Statistically real. Commercially immaterial at most budget levels. Significance and commercial relevance are different questions. Apply both tests before presenting a result to a board.
The internal linking between this post and the Bayesian incrementality approach matters here: a single incrementality test produces a point estimate. Feeding it back into a Bayesian MMM prior is how you turn one clean experiment into a self-correcting measurement system over time.
When incrementality beats attribution and MMM
Each measurement approach earns its place in a specific context. Attribution is appropriate for operational decisions: which keyword to pause, which creative to scale. It is wrong for strategic ones: whether to keep the channel at all, whether a new channel is generating net-new customers or cannibalising existing ones.
MMM answers portfolio-level questions about marginal return across channels and seasons. It requires volume (typically two years of consistent spend history as a floor) and accounts for media interactions and decay in ways no channel-level report can.
Incrementality testing earns its place in four situations: when attribution cannot be trusted because channel overlap is high; when evaluating a new channel before committing budget; when MMM cannot isolate a channel cleanly because the spend is too small or too correlated; and when a board decision requires causal proof of net-new business rather than an attribution story.
The right stack for most organisations is incrementality tests feeding calibration priors into an MMM, with attribution retained for day-to-day campaign decisions. For the maths your platform hides, see our marketing ROI calculators. The full measurement architecture case is in the MMM revival post.
Five-market note
In Singapore, Malaysia, Australia, the US, and Canada, geo holdout design faces different boundary constraints. SG is a city-state: true geo suppression is not possible within the market. Use platform-level conversion lift studies or ghost bids in SG-only campaigns instead of geo holdouts. In AU, state-level suppression works well for direct-response campaigns with low geographic spillover. In Canada and the US, province and state level suppression is standard. MY campaigns can segment by region (Klang Valley vs East Malaysia) for category-level incrementality, though commuter corridor leakage is material in the Klang Valley.
