guide · knowing what you bought

Incrementality testing on a small budget: a worked power table

The power check never sees your budget. It sees your conversions and how much they swing.

Whether a small advertiser can run a valid incrementality test depends on two numbers about its own business: how many conversions the test market produces a day, and how much that count swings from one day to the next. The budget matters only through them. At 5% significance and 80% power, a market averaging 10 conversions a day needs about 70 days of exposure to detect a 15% lift, and one averaging 50 a day needs about 14. The tables below use the formula AdBuyMCP's own power check uses, and every lift in them is a design input: the effect you want the test to be able to see.

A short application: five questions, no card. Approved accounts get the whole loop in their own sandbox, up to a monthly cap. Live spend is open to paid design partners.

By James Rooney Updated

Settings
5% significance, 80% power
10 a day, 15% target
70 days needed
What the check reads
Your conversions and their swing
Length
9 min read
evidenceNo campaign outcomes
On this page

key points

What to remember

These are the points most likely to change your decision, and the full guide follows below.

  1. 01

    The power check reads your test market's daily conversions and how much they vary. Budget matters only through the conversions it adds, so no budget figure is enough on its own.

  2. 02

    Required days rise with the square of the coefficient of variation (the day-to-day swing as a share of the average) and fall with the square of the lift you want to detect. Halving the lift you are chasing quadruples the days.

  3. 03

    At the noise floor for a count, a design needs about 157 extra conversions in the test market to detect a 10% lift and about 78 to detect 20%, whatever the daily volume.

  4. 04

    Multiply those extra conversions by your own incremental cost per conversion to get the media the test market needs. That multiplier is yours, and somebody else's figure cannot stand in for it.

  5. 05

    If your row says hundreds of days, change the design: chase a bigger effect, measure a steadier outcome, or test in a market with more volume.

01

What the power check actually reads

A geo-lift test compares markets you bought with markets you held out. Before the flight, a power check asks whether that comparison could detect an effect of the size you care about. AdBuyMCP runs the check twice: before launch, on London's pre-flight conversions against a fixed 15% target lift, and after the flight, when you ask it to prove a result, on the test market and target you name.

Neither check reads your budget. The pre-launch check accepts a media figure and never uses it. What it reads is the mean and the standard deviation of daily geo-tagged conversions in the test market before the flight, the number of flight days and the lift you want to be able to detect. So the honest answer to "is my budget big enough for a lift test?" is a question back: how many conversions does your test market produce a day, and how much does that number swing?

Source: Meta Open Source, GeoLift walkthrough (open-source docs).

02

The arithmetic and its assumptions

The check is a two-sample comparison of daily means. Take z as the sum of the two critical values, about 2.80 at 5% significance and 80% power; CV as the coefficient of variation of daily conversions, the standard deviation divided by the mean; and L as the lift you want to detect, as a fraction. The days needed are 2 × z² × CV² ÷ L², rounded up, with a minimum of two. Run it the other way and the smallest lift a flight of a given length can detect is CV × √(2 × z² ÷ days).

The formula carries assumptions, and each one makes a real test a little harder than the tables suggest. It treats days as independent, with no allowance for one day's sales predicting the next. It takes the variance only from the test market's pre-period. It uses a normal approximation at the planning stage while the readout uses a Student-t distribution, so very short tests look slightly more feasible here than they are. And it ignores both the noise that well-matched control markets remove and the noise that estimating the pre-period ratio adds.

The lift is always an input. The check never estimates it from spend, and neither should you: it is the smallest effect you would act on, chosen before the test.

03

Days needed, by noise and target lift

Find your coefficient of variation first. Take a few weeks of daily conversions in the market you would test, divide the standard deviation by the mean, and read across to the smallest lift you would act on.

Post-exposure days a geo test needs, at 5% significance and 80% power

Day-to-day swing (CV)10% lift15% lift20% lift30% lift50% lift
0.10167422
0.153616942
0.2063281673
0.301426336166
0.40252112632811
0.50393175994416
0.707703421938631
1.00157069839317563
Computed with the product's power formula when this page is built. Each lift is the effect the design would need to be able to see, chosen in advance. None is a result.
04

The smallest lift a flight can detect

Most small advertisers start from a flight length rather than a target lift. Read across from your noise to the length you can afford, and the cell is the smallest lift that flight could reliably detect. If the effect you expect is smaller than the cell, the test cannot see it.

Smallest detectable lift for a given flight, at 5% significance and 80% power

Day-to-day swing (CV)14 days28 days42 days56 days84 days
0.1010.6%7.5%6.1%5.3%4.3%
0.1515.9%11.2%9.2%7.9%6.5%
0.2021.2%15.0%12.2%10.6%8.6%
0.3031.8%22.5%18.3%15.9%13.0%
0.4042.4%30.0%24.5%21.2%17.3%
0.5052.9%37.4%30.6%26.5%21.6%
Computed with the product's power formula when this page is built. A cell is the smallest effect the design could detect, and says nothing about what a campaign would achieve.
05

From daily conversions to a verdict

If you only know your volume, this table uses the least a count of conversions can swing. Real series swing more, because weekday patterns, promotions and paydays add to it, so treat each row as the best case.

The practical line it draws: a 28-day pilot can detect a 15% lift only if the test market's daily conversions vary by 20% or less, which at this floor means at least 25 conversions a day.

Daily volume in the test market, at the least noise a count can have

Daily conversionsNoise floor (CV)Days to detect 15%Smallest lift in 28 days
20.7134952.9%
50.4514033.5%
100.327023.7%
200.223516.7%
500.141410.6%
1000.1077.5%
Assumes Poisson noise, where the coefficient of variation is 1 divided by the square root of the daily mean. A real series needs at least as many days as its row shows.
06

Turning conversions into a media budget

The check never sees money, so the bridge to a budget is arithmetic you do with your own numbers. The extra conversions a test market must produce are the lift times the daily average times the days. At the noise floor that simplifies to 2 × z² ÷ L whatever your volume: about 157 extra conversions to detect a 10% lift, 105 for 15%, 78 for 20% and 52 for 30%. A noisier series needs more.

Multiply that by your own incremental cost per conversion, the cost of a conversion that would not have happened without the media, and you have the media the test market has to carry during the flight. We do not supply that multiplier. Your channel mix, your offer and your market set it, and a figure borrowed from someone else's business would make the whole calculation look precise and mean nothing.

The spend also has to land where geography is reported, and the control markets must receive none of it. A national channel with no geographic split cannot have its controls verified, so a test built on one can only ever be graded indicative.

Extra conversions needed
The lift × the daily average × the days, or about 2 × z² ÷ L at the noise floor.
Media the test market needs
The extra conversions × your own incremental cost per conversion.
Where it has to land
In markets that report geography, with none of it in the controls.
07

When the table says no

An underpowered test produces a number that says more about which fortnight you ran than about the media, and once that number is in a board pack nobody can tell it from a real one. AdBuyMCP refuses to print a lift from a design that fails the check and returns the days the design would need instead. When your row says no, these are the moves that change it.

Chase a bigger effect
Design for the smallest lift you would act on. If a 10% lift would not change your plan, stop designing for it: the days fall by a factor of four between 10% and 20%.
Pick a steadier outcome
Measure something closer to the media that swings less day to day, such as qualified enquiries, or add up a sparse daily series into weeks.
Concentrate
Put the test in your highest-volume market and hold out markets that behave like it, which raises the daily average the formula works on.
Report what you can support
When no design fits, run the campaign for delivery and observed response, each under its own label, and keep the causal claim for a test that can carry it.

The version of this that matters is the one with your own audience in it. Apply for a sandbox of your own, and once it is approved every channel gets scored against that audience, with nobody else in the room.

Start now

A five-question application, no card.

what this guide does not claim

The limits of this guide

This guide prints no lift that anyone achieved. Every percentage in the tables is a design input, the effect a test would need to be able to see, and every day count is what the product's power formula returns for those inputs at 5% significance and 80% power. The formula's assumptions are stated above, and a real series usually needs more days than its row shows. AdBuyMCP has produced no causal result: nothing has run live, and a geo-lift is graded causal only when every geographic outcome is observed, every delivery row maps to a live line and control exposure is verified. The sandbox cannot meet those conditions by construction, and its deterministic baselines swing far less than a small advertiser's conversions, so sandbox output says nothing about what your business can power. No budget figure on this site, including the entry points on the pricing page, is a claim that a lift test can be powered.

If one of those limits is disqualifying, it is better established now than in week three, and a call establishes it in forty-five minutes.

where this came from

The source of every figure above

Outside figures link to the page that published them, each opened on 24 September 2026. Product facts name the code or data module they are read from.

Published sources

  1. 01Meta Open Source, GeoLift walkthroughOpen-source docs · opened 24 September 2026

Our code and data

  • 01packages/measurement/src/geo-lift.ts in our repo: the power formula, its 5% and 80% defaults and the two-day minimum
  • 02apps/api/src/campaign-readiness.ts: the pre-launch check, its London test market, its fixed 15% target lift and the media figure it accepts and does not read
  • 03apps/api/src/routes/campaigns.ts: the post-flight check, its defaults, the refusal and the claim grading
  • 04The three tables and the conversion arithmetic are computed on this site from that formula when the page is built, so a change to an input moves every cell

This guide is 1,950 words long, counted from its text when the page is built. Where the platform’s README and its code disagree, the code wins.

bring a brief

We'll run all of this on your own audience.

In a 45-minute session, one sentence about your audience becomes seven channel plans in front of you, with the fidelity score, the lawful-basis manifest and the measurement eligibility on screen.

A five-question application, no card.

Bring a real brief to a 45-minute session and we’ll plan it live. · Design-partner phase · a sandbox account is granted on a five-question application, no card

Questions this guide gets asked

Answered in full here, and indexed alongside every other question this site answers at /faq.

How many conversions a day do I need for a geo-lift test?

It depends on the lift you want to detect and how long you can run. At the noise floor, a 28-day test can detect a 15% lift from about 25 conversions a day in the test market, and a smaller effect or a shorter flight needs more. Real series are noisier than the floor, so check your own coefficient of variation against the first table before you plan the flight.

Is there a minimum budget for an incrementality test?

No budget figure answers that on its own, because the power check never reads the budget. What decides it is the extra conversions the test market has to produce, which the tables give you, times your own incremental cost per conversion. Compare that product with what the test market will actually receive, and if it falls short, change the design before you spend.

Why does the sandbox always pass the power check?

Because its baselines are deterministic fixtures scaled to population, and they swing far less than a real advertiser's conversions: in every sandbox market the day-to-day variation is around a tenth of the mean. That exercises the method and says nothing about what your business can power, which is why this guide uses stated assumptions and leaves sandbox output out.