procedure

Run a geo-lift test

The two refusals are steps, not errors.

A geo-lift test runs advertising in some markets, withholds it from comparable others, and compares the outcomes. It is the causal design that needs no login, pixel or match rate, which makes it the only credible option for out-of-home, cinema and television. Most such tests fail before they start, for two reasons: not enough power to detect a plausible effect, or a control group that received some of the campaign. Both failures are steps in this procedure rather than errors in it, and the engine returns each one with the specific thing that would fix it.
steps
6
roughly
Set up in an hour; the test runs for weeks
things needed first
4
before you start

What you need first

  • Enough geographically separable markets — a national campaign in one market cannot be tested this way
  • An outcome you can observe by geography: bookings, leads, point-of-sale, or conversions with a location
  • Pre-campaign history in those markets, so they can be matched on behaviour rather than on demography
  • Enough spend that a plausible effect would be detectable, which is what the power step establishes
the procedure

6 steps

Every step carries the thing that goes wrong at it, in its own block. That is the part worth reading.

  1. step 01

    Nominate a test market

    Choose where the media will run. AdBuyMCP matches candidate control markets against it on the pre-campaign outcome series rather than on demographic similarity.

    what goes wrong here

    Matching on demography is the intuitive choice and the wrong one. Two cities that look alike on paper often behave differently commercially, and it is the behaviour you need to match.

  2. step 02

    Let the power analysis run first

    A two-sample calculation on daily observations runs before any readout, defaulting to 5% significance and 80% power. It returns either a powered verdict with the minimum detectable lift, or an underpowered one.

    what goes wrong here

    This is the step that saves the money. An underpowered test does not produce a weak answer, it produces a random one.

  3. step 03

    If it refuses, read what it needs

    An underpowered design returns the number of days of post-exposure data required and the minimum lift the design could detect, with no lift figure at all. Its own words are that reporting a number from it would be reporting noise.

    what goes wrong here

    If the minimum detectable lift is 30% and a realistic lift is 5%, the test as designed cannot answer your question. The engine returns the sample days that would close the gap instead of a number, and if that is longer than the flight you can fund, the test should not be commissioned.

  4. step 04

    Check the control is genuinely held out

    If every candidate control market received delivery from the campaign, none of them is a holdout and the engine refuses on that basis, naming the contaminated markets.

    what goes wrong here

    A contaminated control is the most common way a lift test quietly measures nothing, because it still produces a confident-looking number. Exclude those markets from targeting for the test window, or nominate a control the campaign does not buy.

  5. step 05

    Run the flight without touching the design

    Let the campaign run. Changing the targeting, the budget split or the market list mid-flight invalidates the comparison.

    what goes wrong here

    The temptation to optimise into the control markets when the test markets are performing is strong and fatal.

  6. step 06

    Read the difference-in-differences result

    A powered design reports the lift with its confidence interval, comparing the change in exposed markets against the change in control markets so that anything affecting both cancels out.

    what goes wrong here

    A result is graded causal only when every geo outcome is observed, every delivery row maps to a live line, and control exposure is verified. Sandbox delivery is labelled modelled by construction, so a sandbox campaign can never reach a causal grade.

what you end up with

Either a lift figure with a confidence interval that can survive being questioned, or a documented refusal naming exactly what would have to change.

45 minutes. Bring a real brief and we compile it live.

Talk it through

Questions

How much spend does a geo-lift test need?

There is no universal figure, and anyone quoting one is guessing. It depends on your baseline variance, your conversion volume, how many independent markets you have and how large an effect you would need to detect. That is precisely what the power analysis computes before the test runs, which is why it runs first. A multi-catchment retailer can often power a test that a single national campaign cannot at any budget.

Has AdBuyMCP produced a causal result for a customer?

No. Nothing has run live, and a geo-lift is graded causal only under conditions sandbox data cannot meet by construction. The engine, the power gate and both refusals are real and tested; the result is not yet evidence about a real campaign, and no lift figure appears anywhere on this site as though it were.