guide · knowing what you bought

Why most incrementality tests are invalid

An underpowered test does not give a weak answer. It gives a random one.

Most incrementality tests fail before the media runs, and for two reasons that are both checkable in advance. The first is power: if the design cannot detect an effect of a plausible size given the spend, the conversion volume and the variance in the data, then any number it produces is noise that happens to look like a result. The second is contamination: if the control markets received some of the campaign, the comparison is exposed against slightly-less-exposed and the measured lift is an underestimate of unknown size. Neither failure announces itself in the output — both produce a confident-looking percentage — which is why the checks have to happen before the money is committed rather than after.
Standard significance
5%
Standard power
80%
Failure one
Underpowered design
Length
6 min read
evidenceNo campaign outcomes
Book a working session
Bring a real brief and everything described here runs against it: the compiled plans, the fidelity scores, and the question of whether the campaign can be proved at all.
Book a working session →
On this page
if you read nothing else

What to remember

The whole guide is below. These are the parts that change a decision.

  1. 01

    Power is the probability a test detects a real effect. A design with low power usually misses genuine effects and its positive findings are as likely to be noise as signal.

  2. 02

    The minimum detectable effect is the number that tells you whether a test is worth running. If it is 30% and a realistic lift is 5%, the experiment as designed cannot answer the question, and the engine returns the sample days it would take to close that gap rather than a lift figure.

  3. 03

    A contaminated control is the most common silent failure. If the campaign reached the control markets, the comparison measures the wrong thing and still returns a number.

  4. 04

    Both checks are arithmetic and both can be done before spending. The reason they usually are not is that the answer is often 'this test cannot work', which nobody wants to hear.

01

What an incrementality test is actually asking

Attribution asks which touchpoints preceded a conversion. Incrementality asks a harder question: would the conversion have happened anyway? The only way to answer it is to compare people who saw the advertising against comparable people who did not, and the entire difficulty is in the word comparable.

This is why retargeting flatters itself so reliably. It reaches people who already visited your site, many of whom were going to convert regardless, and last-click attribution credits it for all of them. An incrementality test is the mechanism that separates the ones it caused from the ones it caught.

For channels where person-level tracking is impossible — cinema, out-of-home, television — a geographic experiment is the only credible causal design available. That makes the quality of the design the whole ball game.

02

Failure one: not enough power to answer

Statistical power is the probability that a test detects a real effect of a given size. The convention is 80% power at 5% significance, and a design that falls short of it will usually miss a genuine effect — and, more dangerously, any positive result it does produce carries little information.

The number that makes this actionable is the minimum detectable effect: the smallest true lift the design could reliably find. If your design's minimum detectable effect is a 30% lift and the realistic effect of your campaign is 5%, no amount of patience fixes it. The test cannot answer your question and should not be commissioned.

The inputs are your baseline conversion volume, the variance in that baseline, the number of independent markets and the duration. All of them are known before the campaign runs, which is why this is a pre-flight check rather than a post-hoc caveat. AdBuyMCP computes it before any readout and refuses to produce a lift figure when the design falls short, returning the required sample days and the minimum detectable lift instead — its own phrasing is that reporting a number from it would be reporting noise.

03

Failure two: the control was not held out

A holdout only works if it is genuinely held out. The common failure is that the campaign reached the control markets — through national inventory, spillover from an adjacent region, or a channel whose targeting is coarser than the test design assumed — so the comparison is between heavily exposed and lightly exposed rather than exposed and unexposed.

The result is an underestimate of the true lift by an unknown amount, presented with the same confidence as a clean result. Nothing in the output flags it. This is why the check has to be structural: verify that no candidate control market received delivery before running the comparison.

The fix, when contamination is found, is specific rather than general: exclude those markets from targeting for the duration of the test window, or nominate a control the campaign does not buy. AdBuyMCP refuses to read out a contaminated design and names the affected markets rather than adjusting silently.

04

Choosing markets on behaviour, not demography

The intuitive way to match a control market to a test market is on demography: similar size, similar income, similar age profile. It is also the wrong way. Two cities that look alike on paper frequently behave differently commercially, because of competitor presence, distribution, local seasonality or a dozen other things nobody is modelling.

The right basis is the pre-campaign outcome series itself. If two markets tracked each other closely on the metric you care about before the campaign, a divergence afterwards is meaningfully attributable. If they did not, no amount of demographic similarity rescues the comparison.

This is also what makes the analysis robust to shared shocks. Difference-in-differences compares the change in the test markets against the change in the controls, so anything affecting both — seasonality, a competitor's national campaign, the weather — cancels. What does not cancel is anything that hit only the test markets, which is the residual risk in any geo experiment and the reason a result should carry a confidence interval rather than a single number.

05

What to ask when you are handed a result

Four questions separate a result worth acting on from one worth ignoring. What was the minimum detectable effect, and is the reported lift comfortably above it? How were the control markets chosen, and on what series were they matched? Was contamination checked, and what did the check find? And what is the confidence interval, rather than the point estimate?

A result presented as a single percentage with no interval and no design detail is not evidence. It may be true, but nothing about how it was presented allows you to tell.

The corollary is worth stating plainly: a vendor who will tell you when a test cannot work is more useful than one who always returns a number. The refusal carries information; the number might not.

If the version of this that matters is the one about your own budget, that is a working session rather than a page.

Talk it through
what this guide does not claim

The limits, in the same size type as the rest

This guide explains why tests fail and what to check. It does not claim AdBuyMCP has produced a causal result for a customer, because it has not — nothing has run live here, and a geo-lift is graded causal only when every geo outcome is observed, every delivery row maps to a live line and control exposure is verified, which sandbox data cannot satisfy by construction. The engine, the power gate and both refusals are real and tested. No lift figure appears anywhere on this site as though it were evidence about a real campaign.

If one of those limits is disqualifying, it is better established now than in week three, and a call establishes it in forty-five minutes.

Talk it through
where this came from

Every figure above, and the file it was read from

Named rather than linked. A URL nobody opened on the day it was attached is a citation in appearance only, so this names the code, the data module or the dated research report instead, and you can go and check.

  • 01The power analysis defaults, the refusal behaviour and the contamination check are read from the geo-lift engine and the campaign routes in the product repository.
  • 02The difference-in-differences readout and the claim grading are from the same source.

1,347words, counted from this page rather than claimed. Where the platform’s README and its code disagree, the code wins.

// bring a brief

Everything above, run against your own audience.

Forty-five minutes. One sentence compiles into seven channel plans in front of you, with the fidelity score, the lawful-basis manifest and the measurement eligibility on screen rather than described.

Book a working session

45 minutes. Bring a real brief and we compile it live. · Design-partner phase · the sandbox needs no card and no credentials

Questions this guide gets asked

Answered in full here, and indexed alongside every other question this site answers at /faq.

Can I run an incrementality test on a small budget?

Sometimes, and the deciding factor is rarely the budget alone. It is the combination of conversion volume, baseline variance and the number of independent markets you can separate. A multi-catchment retailer with twenty locations and steady weekly bookings can often power a test on a modest budget; a national brand with one campaign and lumpy conversions frequently cannot power one at any budget. The power calculation answers it in advance, which is the point of running it first.

Is marketing mix modelling a substitute?

No, though it answers an adjacent question. Modelling estimates each channel's contribution from historical patterns without needing identifiers, which makes it valuable for channels that cannot be tracked. It remains correlational, hungry for history and sensitive to specification — two competent modellers can produce materially different answers from the same data. It is evidence, but it is not the same evidence as a controlled experiment, and presenting a model output as proof of causation is the failure the four-verb model exists to prevent.