Data · Measurement

Incrementality testing: the only ROI proof a board believes

Attribution tells you where the platform saw a user before a conversion. It does not tell you whether the ad caused it. Incrementality testing does. Here is how to design a test, avoid the failure modes, and produce a result a board can act on.

Incrementality testing and geo holdout ROI proof: ink outline of two mirrored bar charts for test and control groups, a segmented horizontal pipe for split experiment flow, a coordinate axis for lift measurement, and a small solid flat orange circle marking the causal signal point.

Bottom line

Attribution tells you which touchpoints preceded a conversion. Incrementality testing tells you whether the ad caused it, by running a controlled experiment with a holdout group that does not see the campaign.

  • Three methods: geo holdouts (suppress a matched geography), conversion lift studies (platform-side user randomisation), ghost bids (phantom auction bids without ad delivery).
  • Design failures kill more tests than statistical ones: match on your business metric, pre-register the outcome, size for power, lock the suppression perimeter.
  • Read results for novelty effects, contamination, and the gap between statistical and commercial significance.
  • Incrementality and MMM are complementary: experiment results feed calibration priors back into the model.

Why attribution is not enough

Every ad platform hands you an attribution report. You spend money, someone buys, the platform takes credit. Run two platforms, and both claim the same sale. Run three, and you are paying for a fiction.

Attribution tells you where the platform saw a user before a conversion. It does not tell you whether the ad caused the conversion. Those are different questions, and conflating them has funded years of wasted spend.

Incrementality testing asks the causal question directly: if you removed this campaign, would the result have happened anyway? The answer comes from a controlled experiment, not from window-matching or weighted probability scores.

This matters more as budgets tighten. A board does not want to hear that your platform dashboard showed a 5x ROAS. They want to know whether removing that channel would change revenue. Incrementality is the method that produces an honest answer. MMM gives you the portfolio view; individual tests give you the isolated signal. They are complementary, not competing. For the relationship between them, see our post on why MMM has returned as a board-level measurement tool.

Three methods dominate. Each has a different set of trade-offs in accuracy, cost, and operational complexity.

The three methods: geo holdouts, conversion lift, and ghost bids

These methods differ in what they suppress, what they measure, and what assumptions they require you to accept.

Geo holdouts suppress advertising in a matched geography while running normally in a test geography. You measure the outcome difference between the two over the test window. The design is simple enough that you control it entirely without platform cooperation. Matched-pair design is the core requirement: you must select geographies that tracked each other closely on the relevant business metric before the test window began.

The holdout window varies by channel. Short-cycle campaigns (paid search) need at least four weeks for the effect to settle. Channels with longer adstock cycles (programmatic, CTV) typically need six to eight weeks before you can trust the delta. The existing leapbuzz practice on geo holdout calibration, documented in our post on cookieless MMM and Bayesian incrementality in APAC, covers channel-specific window guidance in detail.

Conversion lift studies use platform-side randomisation to split audiences into test and holdout cells. Users in the holdout cell are withheld from the campaign; their conversion behavior is compared to the test cell over the study period. Major social and programmatic platforms offer this natively, typically as a feature available to accounts above a volume threshold. The advantage is user-level precision across a single channel. The limitation is channel isolation: the holdout users may be reached by the same campaign on a different platform, contaminating the control.

Ghost bids are the cookieless answer to conversion lift. Instead of withholding users from the auction, the system submits phantom bids for users who would have been targeted, then tracks their downstream behavior without showing them the ad. The conversion delta between ghost-bid users and served users estimates the incremental effect. The method requires technical implementation at the platform or DSP level and is not universally available. It has grown in relevance as identity signals have narrowed.

Which method to start with

Geo holdouts require no platform relationship and can be run on any channel, including offline. Start there. Graduate to conversion lift when you need user-level precision on a single social channel and the platform relationship supports it. Use ghost bids when cookie-based user matching is unreliable in your market.

Designing a first test

Most incrementality tests fail at the design stage, not the measurement stage. Four decisions determine the validity of everything that follows.

Select the outcome metric before the test starts. Pick one. Revenue, new customers, store visits, app installs. Pre-registration matters because outcome switching after the test is a form of data manipulation that looks like analysis. If you are unsure which metric the business cares about, resolve that before designing the test.

Match geographies on the business metric, not population. Two cities of equal size are not a valid matched pair if one has twice the organic purchase rate. Match on the conversion rate or revenue-per-population of your specific outcome over a historical baseline of at least twelve weeks. Seasonal matching matters too: avoid testing windows that span major shopping events unless both geographies experience them identically.

Size the test for statistical power before running it. The calculator below gives you a quick power estimate. The practical implication: small budgets and low conversion rates require very large test populations or very long test windows to detect moderate lifts at acceptable confidence. If your test is underpowered, you will either declare a false negative (the campaign worked but the test was too small to show it) or make an irreversible budget decision on noise. A minimum detectable effect of 10 percent lift on a 2 percent baseline conversion rate requires far more traffic than most teams expect.

Lock down the suppression perimeter. For geo holdouts, map which campaigns run in the holdout geography before the test window. Any spend leaking into the holdout breaks the design. For conversion lift, confirm whether the holdout cell receives the same campaign on other platforms in your media mix. A user in your Meta holdout who is being served the same message on YouTube is not a clean control.

These four decisions made in advance will do more for result quality than any statistical sophistication applied after the fact.

Uplift significance calculator

Use this tool to estimate whether an observed lift is statistically meaningful or within noise range. Enter the visitor counts and conversions from your test and control groups. The calculator runs a two-proportion z-test and gives you a plain-language read on the result.

Incrementality uplift calculator

Planning model. Enter test and control group data to estimate lift significance.

Control conversion rate
Test conversion rate
Absolute lift
Relative lift
Z-score
Approximate confidence

Planning model only. Two-proportion z-test assumes independent groups, simple random assignment, and binary outcomes. It does not account for clustering in geo holdouts, covariate imbalance, or sequential testing inflation. Use it to size tests and sense-check results, not as a final statistical authority. For geo holdouts in particular, consult a statistician before acting on results near the significance boundary.

Reading results without fooling yourself

A positive lift number is not a pass. Three failure modes regularly inflate incrementality results in ways that look clean until you go looking for them.

Novelty effects. Users respond differently to a new campaign in its first week or two. Response rates are inflated by curiosity or freshness, then settle. If your test window is four weeks or shorter, the average lift may be driven by the opening burst rather than steady-state performance. A longer test window or a pre-analysis period that excludes the first week of exposure helps here. The effect is more pronounced for brand new creative than for refreshed campaigns.

Contamination. In geo holdouts, contamination enters when users cross your geographic boundary. A campaign suppressed in city B reaches users who drive, commute, or shop in city A. A user in your Meta holdout cell who sees the same campaign on a different platform is contaminated. Neither is a design failure you can always prevent, but you need to estimate its likely magnitude before reporting the result. High commuter corridors between test and holdout geographies are a known risk; select boundaries that align with commercial rather than residential geography where possible.

Selection bias in conversion lift cells. Platform-side holdout randomisation is not always truly random at the user level. Platforms assign holdout cells based on their own optimisation logic, which can systematically put lower-intent or harder-to-reach users into the control cell. If you see an unusually high conversion rate in your holdout group, that is a signal of potential cell imbalance, not evidence that your campaign hurt performance. Ask the platform for cell composition diagnostics where available.

One other read that trips teams up: a statistically significant lift at a small absolute size. A campaign that generates a 15 percent relative lift on a 0.3 percent baseline conversion rate has moved from 0.3 to 0.345 percent. Statistically real. Commercially immaterial at most budget levels. Significance and commercial relevance are different questions. Apply both tests before presenting a result to a board.

The internal linking between this post and the Bayesian incrementality approach matters here: a single incrementality test produces a point estimate. Feeding it back into a Bayesian MMM prior is how you turn one clean experiment into a self-correcting measurement system over time.

When incrementality beats attribution and MMM

Each measurement approach earns its place in a specific context. Attribution is appropriate for operational decisions: which keyword to pause, which creative to scale. It is wrong for strategic ones: whether to keep the channel at all, whether a new channel is generating net-new customers or cannibalising existing ones.

MMM answers portfolio-level questions about marginal return across channels and seasons. It requires volume (typically two years of consistent spend history as a floor) and accounts for media interactions and decay in ways no channel-level report can.

Incrementality testing earns its place in four situations: when attribution cannot be trusted because channel overlap is high; when evaluating a new channel before committing budget; when MMM cannot isolate a channel cleanly because the spend is too small or too correlated; and when a board decision requires causal proof of net-new business rather than an attribution story.

The right stack for most organisations is incrementality tests feeding calibration priors into an MMM, with attribution retained for day-to-day campaign decisions. For the maths your platform hides, see our marketing ROI calculators. The full measurement architecture case is in the MMM revival post.

Five-market note

In Singapore, Malaysia, Australia, the US, and Canada, geo holdout design faces different boundary constraints. SG is a city-state: true geo suppression is not possible within the market. Use platform-level conversion lift studies or ghost bids in SG-only campaigns instead of geo holdouts. In AU, state-level suppression works well for direct-response campaigns with low geographic spillover. In Canada and the US, province and state level suppression is standard. MY campaigns can segment by region (Klang Valley vs East Malaysia) for category-level incrementality, though commuter corridor leakage is material in the Klang Valley.

Frequently asked questions

What is incrementality testing in marketing?

Incrementality testing is a controlled experiment that measures whether an advertising campaign caused a conversion, not just preceded it. You split audiences or geographies into a test group that sees the campaign and a holdout group that does not. The difference in conversion rates between the two groups is the incremental effect of the campaign. Unlike attribution, which assigns credit based on touchpoint proximity, incrementality produces a causal estimate: how many conversions would not have happened without the spend.

What is the difference between a geo holdout and a conversion lift study?

A geo holdout suppresses advertising in a matched geography and runs normally in a test geography. The two geographies are compared over a defined window. You control the design without platform cooperation and can test any channel, including offline. A conversion lift study uses platform-side user-level randomisation to assign individuals to test and holdout cells. It is more precise for a single channel because it tracks specific users, but it requires platform access and is vulnerable to cross-platform contamination if the holdout users are reached by the same campaign elsewhere.

How long should an incrementality test run?

Window length depends on the channel's adstock cycle and your baseline conversion volume. Paid search campaigns can settle in four weeks. Paid social typically needs six weeks for the novelty effect to decay and the steady-state signal to emerge. Programmatic, CTV, and out-of-home campaigns commonly need eight weeks due to longer adstock half-lives. Short tests run the risk of being dominated by novelty-effect inflation. Run the test until you have adequate conversion volume in both groups to detect your minimum detectable effect at the confidence level you need for a business decision.

What is a ghost bid in incrementality testing?

A ghost bid is a phantom bid submitted in a real-time auction for a user who would have been targeted, but without serving the ad. The system tracks the downstream conversion behaviour of ghost-bid users and compares it to users who were actually served. The delta estimates the incremental effect of the ad exposure. Ghost bids do not require cookie-based user matching in the holdout, which makes them useful when identity signals are narrow or consent limitations constrain standard holdout randomisation. The method requires implementation at the DSP or platform level and is not universally available.

How do you select matched geographies for a geo holdout?

Match on the specific business metric you are testing, not on population. Two cities of equal size are not a valid pair if one has a significantly higher organic purchase rate. Use at least twelve weeks of pre-test history to establish baseline parity on your outcome metric. Avoid boundaries with high commuter flow between geographies, which creates contamination risk. Exclude geographies with scheduled local events (major sports, festivals) that fall inside the test window. The baseline correlation between test and holdout on the outcome metric is your quality check: the higher it is, the more confident you can be that post-test differences are caused by the campaign rather than pre-existing divergence.

What is statistical power and why does it matter for incrementality tests?

Statistical power is the probability of detecting a real effect when one exists. An underpowered test produces a false negative: the campaign was actually incremental, but the test was too small or too short to show it. Before running a test, calculate the sample size needed to detect your minimum detectable effect at your target confidence level, given your baseline conversion rate. Low baseline conversion rates require very large populations to detect moderate lifts. If you cannot reach adequate power within your budget or timeline, the test result will be ambiguous regardless of what the numbers show.

How do novelty effects distort incrementality test results?

Novelty effects occur when users respond to a new campaign with higher-than-normal engagement in the first one to two weeks, then revert to steady-state behaviour. If your test window is short, the average lift will be inflated by this opening burst and will not represent sustainable campaign performance. The distortion is largest for new creative or channels that audiences have not seen before. Mitigation options include extending the test window to capture post-novelty behaviour, excluding the first week of data from the analysis period, or running a pre-period to measure baseline response before the full campaign exposure begins.

When should I use incrementality testing instead of MMM?

MMM gives you the portfolio view across channels, seasons, and spend levels. Incrementality testing gives you the isolated causal signal for a specific channel or campaign. Use incrementality when you need a binary answer on a single channel (is this working or not?), when you are evaluating a new channel before committing budget to the MMM training window, or when the MMM cannot isolate a channel because the spend is too small or too correlated with other activity. The two methods are complementary: incrementality test results feed calibration priors back into a Bayesian MMM, producing a more accurate model over time.

Can incrementality testing work in Singapore given it is a single city-state?

Geographic suppression within Singapore is not practical because the market is a single contiguous city-state with no meaningful regional boundary that would prevent ad exposure from crossing the holdout zone. For Singapore-only campaigns, use platform-level conversion lift studies (user-level holdout randomisation within the platform) or ghost bids as the primary incrementality method. In multi-market plans covering SG alongside AU, MY, US, or CA, you can use country-level geo holdouts to test channels at market level, running the campaign in one country while suppressing it in another over a matched window, though market differences reduce the clean-pair assumption.

Related

Work with leapbuzz

Need a measurement architecture your board will actually believe?

leapbuzz designs incrementality testing programmes and measurement stacks for marketing teams across Singapore, Malaysia, Australia, the US, and Canada. From first geo holdout to a calibrated MMM, we build the causal proof, not the attribution story.

Talk to us