guide · knowing what you bought

Designing a geo-lift test that can actually be powered

The design decides the answer. Run the power calculation before the flight, not after it.

A geo-lift test measures incrementality by holding out matched markets: you buy some places and deliberately do not buy others, and the difference between them after the campaign starts, adjusted for how they differed before it, is the lift. It is the one causal design in media that works without a login, a pixel or a match rate, which is why it suits out-of-home, cinema and television as well as it suits digital. It is also the design most often run in a state where it could never have detected anything, because the sample-size arithmetic was not done first. The number of days a test needs rises with the square of the baseline volatility and falls with the square of the lift you are trying to detect, so a thin, noisy conversion series cannot be rescued by running longer at the margin. AdBuyMCP runs that calculation before the test and returns the required sample days instead of a lift figure when the design cannot support one.
Power gate
α 0.05 · 80% power
Markets in the table
12 UK conurbations
Matching weight
70% correlation, 30% size
Length
10 min read
evidenceNo campaign outcomes
Book a working session
Bring a real brief and everything described here runs against it: the compiled plans, the fidelity scores, and the question of whether the campaign can be proved at all.
Book a working session →
On this page
if you read nothing else

What to remember

The whole guide is below. These are the parts that change a decision.

  1. 01

    Run the power calculation before the flight. It is the only moment when the answer is still actionable, because the fixes are all design changes.

  2. 02

    Required days scale with the square of your baseline volatility and inversely with the square of the lift you want to detect. Halving the lift you are chasing quadruples the test.

  3. 03

    A control market that received delivery is not a control. Exclude the test markets from the candidate list before matching, not after.

  4. 04

    Match on the pre-period series, not on intuition about which cities are similar. Correlation on date-joined daily data does the work.

  5. 05

    An underpowered test does not produce a weak result. It produces a random one, and a random number in a board pack is indistinguishable from a real one.

01

What the design is, in one paragraph

Pick markets that behaved alike before the campaign. Buy media in some of them and deliberately buy none in the others. After the flight, scale the control markets by the ratio the two groups sat at before the media started, and treat that scaled series as the counterfactual: what the exposed markets would have done anyway. The gap between what they actually did and that counterfactual is the estimate, and it is a genuine causal estimate because the holdout was chosen by you rather than selected by exposure.

That last clause is the whole reason this design is worth the trouble. Almost every other measurement in media compares people who saw the ad with people who did not, and the two groups differ in the one way that matters: one of them was targeted. A geo holdout is decided before delivery, by the buyer, on a geography rather than on a person, which is why it survives the collapse of identity-based measurement and why it works on channels with no user-level data at all.

02

The arithmetic that decides whether it is possible at all

The sample-size calculation for a two-sample comparison of daily observations is short enough to write down: the required number of days is roughly two times the squared sum of the two critical values, times the variance of your daily baseline, divided by the square of the absolute change you are trying to detect. At the conventional settings — a 5% significance level and 80% power — that sum of critical values is a shade over 2.8.

Two consequences follow, and both of them are counter-intuitive to anyone who has planned a flight by budget. Volatility enters squared, so a conversion series that swings wildly day to day needs dramatically more days than a smooth one at the same volume. And the lift you are chasing enters squared in the denominator, so a design that can detect a 20% lift in three weeks needs roughly twelve weeks to detect a 10% one. Chasing a small effect is expensive in a way that is not obvious until the calculation is done.

The reason to run it before the flight is that every fix is a design change. Once the media has run, the volatility is what it was and the days are what they were, and the only remaining move is to report a number the design cannot support. AdBuyMCP therefore runs the calculation at planning time and returns one of two things: a powered verdict with the smallest lift the design could detect, or an underpowered verdict with the number of days it would actually need and no lift figure at all.

03

Choosing the markets, and why intuition is the wrong tool

The candidate list is twelve UK conurbations with their populations and broadcast regions: Greater London, Greater Manchester, the West Midlands, West Yorkshire, Greater Glasgow, Liverpool City Region, Sheffield City Region, Tyneside, Greater Nottingham, Bristol, Edinburgh and Cardiff. That set is large enough to give a national campaign a real holdout and small enough that each market carries meaningful weight.

Matching ranks the candidates on two things: how strongly each one's pre-period conversion series correlates with the test market's, weighted at 70%, and how similar the two populations are, weighted at 30%. The top three become the control group. Correlation does most of the work because the design's assumption is about parallel trends rather than about equal levels: a control market half the size is fine if it moves in step.

One implementation detail is worth stealing whatever tool you use. The two series are joined on date before they are correlated, not aligned by position in the array. Real conversion series have gaps — a day with no conversions often produces no row at all — and index-aligning two series with different gaps pairs Tuesday with Thursday and can invert the sign of the correlation. A candidate sharing fewer than two dates with the test market cannot be correlated and is skipped rather than scored at zero.

04

The control you have to protect

The most common way a lift test quietly measures nothing is contamination: the control markets received some of the media. It happens by accident more often than by design — a targeting fallback widening to national, a programmatic line delivering outside the intended polygons, a search campaign with no geographic restriction at all — and it biases the result towards zero, which makes a working campaign look ineffective.

The defence is procedural rather than statistical. Before matching runs, every market the campaign actually delivered into is excluded from the candidate list. If that leaves nothing, the honest response is to refuse: AdBuyMCP returns a "no clean control" refusal naming the contaminated markets, rather than matching against a second test cell and calling it a holdout.

There is a second, quieter failure. Where the buy reports no geography at all — which is normal on several rails — you cannot verify that the control markets were clean, only assume it. Control integrity is then reported as unverified rather than passed, along with the count of delivery rows carrying no geography. A check that cannot see the delivery has not passed; it has not run.

05

Reading the result without over-claiming it

The readout is a difference-in-differences with pre-period ratio scaling. The control series is multiplied by the ratio the two groups sat at before exposure, that scaled series is the counterfactual, and the daily deviations after exposure give the point estimate, the confidence interval and the p-value.

Two details in that calculation are worth insisting on wherever you have it run. Inference should be Student-t rather than normal, because post-exposure windows in media are short and the normal approximation understates the tails badly at two weeks. And the standard error has to carry both sources of uncertainty: the noise in the post-period, and the fact that the scaling ratio is itself an estimate from a finite pre-period. Treating that ratio as a known constant is how a lucky fortnight before the campaign turns into a tight, spuriously significant lift.

Then there is the grade, which is separate from the statistics. Passing the power gate is necessary and not sufficient. A result is only allowed to call itself causal when every geographic outcome is observed, every delivery row used in the control check maps to a line that executed live, and control exposure is verified. Where inputs are modelled or the delivery is sandbox, the result is graded a modelled demonstration; where the evidence is mixed or exposure could not be checked, it is graded indicative. The grade travels with the number.

06

Four things that actually fix an underpowered design

When the calculation says no, there are four real moves and one fake one. The fake one is running it anyway and reporting whatever comes back.

Extend the flight
The direct fix, and the calculation tells you exactly how far. Take the required sample days at face value rather than splitting the difference.
Concentrate the spend
Fewer test markets with more weight each. Concentration matters more than total budget, because the exposed series has to move against its own noise.
Accept a larger detectable lift
Ask whether a 25% effect is what you would act on anyway. Designing for the smallest interesting effect rather than the smallest imaginable one saves weeks.
Reduce the baseline noise
Measure a less volatile outcome closer to the media — qualified enquiries rather than closed revenue — or aggregate to weeks where the daily series is dominated by zero days.
07

What to do when the answer is still no

Some campaigns cannot be proved causally, and the number of them is larger than the incrementality industry likes to admit. A national buy with no geographic dimension in its reporting has nothing to hold out. A thin conversion series at a small budget cannot clear the arithmetic at any sensible flight length. In both cases the useful response is to say so at planning time and offer what the campaign can actually evidence: delivery, observed response on your own properties, and modelled reach and frequency, each under its own label.

That is a smaller claim and it is the one the campaign can carry. It is also a considerably less pleasant conversation to have at the point of sale than at the point of reporting, and a considerably more useful one, because the buyer still has the option of changing the design.

If you take one habit from this guide, take this one: ask whether the campaign can be proved before you fund it, in the same meeting where you agree the budget. The answer takes about ten minutes to compute and it changes what you buy.

If the version of this that matters is the one about your own budget, that is a working session rather than a page.

Talk it through
what this guide does not claim

The limits, in the same size type as the rest

AdBuyMCP has not produced a causal result for anyone. Nothing has run live, and a geo-lift is graded causal only when every geographic outcome is observed, every delivery row maps to a live line and control exposure is verified — conditions the deterministic sandbox cannot meet by construction, because sandbox delivery rows are labelled modelled. Any lift figure visible in a demonstration is injected sandbox data carrying a modelled-demonstration grade, and this guide prints no lift percentage for that reason. The engine, the power gate, the market matching and the refusals are real, tested code; the results are not yet evidence about a real campaign. The baseline series in the sandbox are deterministic fixtures generated from population, a weekly pattern and seeded noise, so they exercise the method rather than describing any real market's conversion behaviour.

If one of those limits is disqualifying, it is better established now than in week three, and a call establishes it in forty-five minutes.

Talk it through
where this came from

Every figure above, and the file it was read from

Named rather than linked. A URL nobody opened on the day it was attached is a citation in appearance only, so this names the code, the data module or the dated research report instead, and you can go and check.

  • 01packages/measurement/src/geo-lift.ts — the twelve-market table, the 70/30 matching heuristic, the power calculation and its α and power defaults, the diff-in-diff with pre-period ratio scaling, the Student-t inference and the Welch–Satterthwaite degrees of freedom
  • 02apps/api/src/routes/campaigns.ts — the prove-it payload, the refusal shapes and the claim grading
  • 03src/data/measurement.ts on this site — the four verbs, the provenance labels and the two published refusals

2,154words, counted from this page rather than claimed. Where the platform’s README and its code disagree, the code wins.

// bring a brief

Everything above, run against your own audience.

Forty-five minutes. One sentence compiles into seven channel plans in front of you, with the fidelity score, the lawful-basis manifest and the measurement eligibility on screen rather than described.

Book a working session

45 minutes. Bring a real brief and we compile it live. · Design-partner phase · the sandbox needs no card and no credentials

Questions this guide gets asked

Answered in full here, and indexed alongside every other question this site answers at /faq.

How many markets do I need for a geo-lift test?

More than one on each side, and enough total volume that the exposed series can move against its own daily noise. Our matcher takes the top three controls per test market, and the candidate list is twelve UK conurbations. The number that actually decides feasibility is not the market count but the required sample days, which comes out of the power calculation on your baseline mean and standard deviation. Two exposed markets with high volume and well-matched controls beat six thin ones.

Can I run a geo-lift test on a national campaign?

Only if you are willing to withhold media somewhere, which is what makes it a test rather than a report. A campaign that buys everywhere has no holdout, and the honest answer is that its incrementality cannot be measured this way. The alternative designs — modelled attribution, marketing-mix modelling — produce estimates rather than causal evidence, and should carry a different label when they do.

Why refuse to report a number when the test is underpowered?

Because the number is not a weak result, it is a random one. An underpowered design produces estimates that swing across a wide range from noise alone, so the figure it returns says more about which fortnight you happened to run than about the media. Once that figure is in a slide it is indistinguishable from a measured one, and it will be defended in a budget meeting six months later. Returning the required sample days instead tells you what would have to change.