Run a geo-lift test
The two refusals are steps, not errors.
- steps
- 6
- roughly
- Set up in an hour; the test runs for weeks
- things needed first
- 4
What you need first
- Enough geographically separable markets — a national campaign in one market cannot be tested this way
- An outcome you can observe by geography: bookings, leads, point-of-sale, or conversions with a location
- Pre-campaign history in those markets, so they can be matched on behaviour rather than on demography
- Enough spend that a plausible effect would be detectable, which is what the power step establishes
6 steps
Every step carries the thing that goes wrong at it, in its own block. That is the part worth reading.
- step 01
Nominate a test market
Choose where the media will run. AdBuyMCP matches candidate control markets against it on the pre-campaign outcome series rather than on demographic similarity.
what goes wrong hereMatching on demography is the intuitive choice and the wrong one. Two cities that look alike on paper often behave differently commercially, and it is the behaviour you need to match.
- step 02
Let the power analysis run first
A two-sample calculation on daily observations runs before any readout, defaulting to 5% significance and 80% power. It returns either a powered verdict with the minimum detectable lift, or an underpowered one.
what goes wrong hereThis is the step that saves the money. An underpowered test does not produce a weak answer, it produces a random one.
- step 03
If it refuses, read what it needs
An underpowered design returns the number of days of post-exposure data required and the minimum lift the design could detect, with no lift figure at all. Its own words are that reporting a number from it would be reporting noise.
what goes wrong hereIf the minimum detectable lift is 30% and a realistic lift is 5%, the test as designed cannot answer your question. The engine returns the sample days that would close the gap instead of a number, and if that is longer than the flight you can fund, the test should not be commissioned.
- step 04
Check the control is genuinely held out
If every candidate control market received delivery from the campaign, none of them is a holdout and the engine refuses on that basis, naming the contaminated markets.
what goes wrong hereA contaminated control is the most common way a lift test quietly measures nothing, because it still produces a confident-looking number. Exclude those markets from targeting for the test window, or nominate a control the campaign does not buy.
- step 05
Run the flight without touching the design
Let the campaign run. Changing the targeting, the budget split or the market list mid-flight invalidates the comparison.
what goes wrong hereThe temptation to optimise into the control markets when the test markets are performing is strong and fatal.
- step 06
Read the difference-in-differences result
A powered design reports the lift with its confidence interval, comparing the change in exposed markets against the change in control markets so that anything affecting both cancels out.
what goes wrong hereA result is graded causal only when every geo outcome is observed, every delivery row maps to a live line, and control exposure is verified. Sandbox delivery is labelled modelled by construction, so a sandbox campaign can never reach a causal grade.
Either a lift figure with a confidence interval that can survive being questioned, or a documented refusal naming exactly what would have to change.
What else you might need to do
45 minutes. Bring a real brief and we compile it live.
Talk it throughQuestions
How much spend does a geo-lift test need?
There is no universal figure, and anyone quoting one is guessing. It depends on your baseline variance, your conversion volume, how many independent markets you have and how large an effect you would need to detect. That is precisely what the power analysis computes before the test runs, which is why it runs first. A multi-catchment retailer can often power a test that a single national campaign cannot at any budget.
Has AdBuyMCP produced a causal result for a customer?
No. Nothing has run live, and a geo-lift is graded causal only under conditions sandbox data cannot meet by construction. The engine, the power gate and both refusals are real and tested; the result is not yet evidence about a real campaign, and no lift figure appears anywhere on this site as though it were.