Measurement

Four verbs, and the refusal that keeps them apart.

Four verbs — delivered, responded, caused, remembered — kept visibly apart, because the entire value of a measurement claim is in which of the four it belongs to. Three are built, the fourth is a placeholder, and no campaign has yet produced a causal result because nothing has run live. What this page sells is the refusal, not a lift number.
Verbs published
4
Built
3 of 4
Power gate
α 0.05 · 80%
Causal results so far
None
the caveat1 of 4 not built
Book a working session
Bring a real campaign and we will run the eligibility question on it: whether your spend, geography and conversion volume can support a causal claim at all.
Book a working session →
The model

Four questions, four different kinds of answer

Almost every argument about advertising measurement is really an argument about which of these four questions a number answers. Keeping them in separate columns is not fastidiousness; it is the only way a delivery report and an incrementality result can sit on the same page without one borrowing authority from the other.

01

Delivered

Did the media run?

Built
What backs it

Normalised delivery rows from each provider, with the additive provider event counts kept side by side rather than merged. Reach and frequency are modelled against the persona universe using Sainsbury deduplication, and anything modelled is labelled modelled.

How far the claim goes

Ratios and cost averages stay dated observations and are never summed across days. A blended CPM computed by adding two days of averages is a made-up number, and the payload refuses to produce one.

02

Responded

Did anyone do something afterwards?

Built
What backs it

Outcomes observed on your own property — pageviews, conversions and their value — through Plausible or GA4, matched on the campaign's immutable UTM key, plus promo codes, vanity URLs and pixel events. Podscribe covers audio attribution.

How far the claim goes

This is correlation and the payload says so in its own note. Two exclusions are enforced rather than advised: the synthetic geo-panel baseline series never inflates response counts, and nothing dated before the flight start can be counted as a response to it.

03

Caused

Would it have happened anyway?

Built
What backs it

A matched-market geo-lift test over UK conurbations with a difference-in-differences readout, run from one button. Before it runs, a two-sample power calculation on daily observations decides whether the test can answer the question at all, at α = 0.05 and 80% power.

How far the claim goes

Two limits, and the second is absolute today. If the design is underpowered the engine returns an underpowered verdict and the required sample days rather than a lift number — that refusal is the point of it. And a result is only graded causal when every geo outcome is observed, every delivery row maps to a live line, and control exposure is verified. Sandbox delivery rows are labelled modelled by construction, so no sandbox campaign can reach a causal grade, and none has: the seeded demonstration asserts its own claim as a modelled demonstration with an injected lift, and is not evidence about anything.

04

Remembered

Did it change what people think?

Placeholder
What backs it

Brand-lift study through Cint: an exposed-versus-control panel measuring recall and consideration. It is the P0 measurement rail for the fourth verb.

How far the claim goes

Nothing here is a substitute for the other three. Brand lift tells you about attitude, not about revenue.

Why it is here anyway

Not built. The API returns this verb with the status "placeholder" and the note that the study has not been commissioned in this build, and the Cint adapter is explicitly non-transacting pending a contract. It is on this page because a four-verb model that quietly ships three verbs is exactly the kind of rounding we exist not to do.

3 of 4 verbs are built. The rememberedverb is on this page at the same size as the other three because the API returns it at the same level of the payload, with the status “placeholder” and a note saying the study has not been commissioned.

If you want to know which of the four your budget can actually reach, that is a question with an answer, and it takes about ten minutes of a call.

Talk it through
Provenance

Four labels, and the rule that stops them drifting

Every number the platform reports carries one of these four labels, and the labels are what keep the verbs distinguishable once a report is exported into a slide. The rule underneath them is blunt: no label may advance on the strength of sandbox or API-shaped data.

01

Provider-reported

The seller's own delivery numbers, as they were returned, in their native dimensions and currency.

02

Observed

Something that actually happened on a property you own and can audit, matched to the campaign by an immutable key. Correlated with the media, not attributed to it.

03

Modelled

An estimate produced by a formula rather than a count — deduplicated reach and frequency being the common case. Always carries the word.

04

Powered causal

The only label that supports the word "caused", and it is only issued by a geo-lift test that passed its own power analysis before it ran.

How the label is decided

A delivery row is graded observed only when the exact persisted campaign line it belongs to executed live. Everything else is modelled, including every row the sandbox produces, and a set that mixes the two is reported as mixed rather than rounded up to observed. That is a comparison in code rather than a statement of intent, which is why a demonstration cannot quietly become a claim between the product and the slide.

Causation

Prove It

One button that runs a matched-market geo-lift test, and refuses to run it when the answer would not mean anything.

  1. 01

    Match the markets

    Candidate UK conurbations are paired on their pre-campaign outcome series so that exposed and held-out markets were behaving the same way before the media started.
  2. 02

    Run the power analysis first

    A two-sample calculation on daily observations, defaulting to α = 0.05 and 80% power, returns either a powered verdict with the minimum detectable lift, or an underpowered verdict with the number of sample days the design would actually need.
  3. 03

    Refuse, or read out

    An underpowered design returns no lift figure at all. A powered one runs difference-in-differences against the held-out markets and reports the lift with its confidence interval.

One button, two outcomes

The power analysis runs before the test, so the design decides which of these two responses comes back. Both are returned, persisted against the campaign and rendered in the product. The refusal is a result with a shape, not an error and not an empty state.

Powered

The design can detect a lift of the size you asked about, so difference-in-differences runs against the held-out markets.

status
completed
power.verdict
powered — the design cleared α = 0.05 at 80% before the test ran
lift.liftPct
the point estimate, with lift.ci95 either side of it
lift.pValue
two-sided Student-t, on Welch–Satterthwaite degrees of freedom
controlMarkets
the matched holdouts, each with its pre-period correlation
controlIntegrity
verified, or the count of delivery rows carrying no geo that stopped it being verified
claim
causal, modelled demonstration, or indicative

Underpowered

The design cannot detect a lift of that size, so the engine returns what the design would need and stops.

status
refused
reason
underpowered
power.requiredSampleDays
the post-exposure days this design would actually need
power.minDetectableLiftPct
the smallest lift this design could detect as it stands
note
a sentence saying that reporting a number from it would be reporting noise
lift
no lift figure at all

Field names as the prove-it endpoint returns them. No lift percentage appears on this page in either column, deliberately: the only figures we have produced so far are deterministic sandbox output, and a sandbox number printed beside a real field name is the thing this whole section argues against.

The refusal is the feature. An incrementality number from an underpowered test is not a weak result, it is a random one, and the reason this engine will not produce one is that a random number presented in a board pack is indistinguishable from a real one.
Why the engine will not return a number it cannot support
The other refusal

A control market that received delivery is not a control. Before matching runs, every market this campaign actually delivered into is excluded from the candidate list, and if that leaves nothing the engine refuses with “no clean control” and names the contaminated markets rather than quietly matching against a second test cell. Where the buy reports no geography at all, control integrity is reported as unverified rather than assumed, because a check that cannot see the delivery cannot pass it.

And what the finished test is allowed to claim

Passing the power gate is necessary and not sufficient. Every completed test is graded on the evidence that fed it, and the grade travels with the number.

Causal
Every geo outcome is observed, every delivery row used in the control check maps to a line that executed live, and control-market exposure is verified. Nothing has met this yet.
Modelled demonstration
The controls check out but at least one input rail is modelled or sandbox delivery. This is the grade every seeded demonstration carries, and it is a demonstration of the method rather than evidence about a campaign.
Indicative
The inputs mix observed and modelled evidence, or no delivery row could be used to verify exposure, or control exposure could not be fully checked. The estimate stands but it cannot support a causal or incremental claim.
// before you fund it

Find out whether your campaign can be proved.

Bring a real brief to a working session. We compile it live, and we tell you which of the four verbs your spend, your geography and your conversion volume can actually support — including when the answer is that a causal claim is out of reach at this budget.

Book a working session

45 minutes. Bring a real brief and we compile it live. · deterministic sandbox · no credentials and no card

The mechanism

One ledger, four keys, and no identity joins inside it

Exposures across all seven channels write into one ledger, keyed on whichever of four identifiers the channel can supply. That single ledger is the mechanism behind cross-channel frequency deduplication and sequential journeys — and it is worth being exact about its status, because the mechanism and the data are different things.

Status, stated exactly

The four-key spine is designed, schema-backed and unit-tested, and today it is fed only by the deterministic sandbox. There is no inbound ingest route for viewshed panels, CTV logs or IP logs, and no shipped code path writes a hashed advertising id, a EUID or a postcode sector from live data — only a household key from the sandbox graph. The refusal is deliberate: a live report proves aggregate delivery, not household exposure, so it must never mint identities that could later cross a vendor boundary. Treat the journey explorer as a working mechanism awaiting live input.

The four keys, in precedence order

An exposure is written against the best key its channel can supply, and the dedup key is taken in this order. Key spaces are prefixed so they can never collide, and the grade travels with the event: a household key or a hashed advertising id grades household, a consent-chained id grades person, and a postcode sector grades geo-cohort.

01
householdKey
An opaque household key issued by the key-translation service. The ledger never resolves it back to anything.
02
maidHash
A hashed mobile advertising id. Held for at most 90 days, then persistently stripped from the event, which re-grades the event rather than deleting it.
03
euid
The consent-chained European Unified ID, where one exists.
04
postcodeSector
The coarsest key and the one that always works. Cinema and DOOH usually land here.
Why four and not one

Multi-key by design, because MAIDs are a declining asset. A spine that only worked on mobile advertising ids would degrade into uselessness on a schedule set by Apple and Google; this one loses precision a step at a time and says which step it is on.

Where identity is allowed to happen

The ledger itself never performs an identity join. Resolution between key spaces happens in one place — the key-translation service — which is the only module we have that touches identity, and it is the module the 90-day rule is enforced in. Keeping that boundary in one file is what makes it auditable.

If the journey explorer is the thing you want, ask about it on a call: it runs today, and today it runs on sandbox exposures.

Talk it through
Before the money moves

Ask whether it can be proved before you fund it

Whether your campaign can support a causal claim is a property of your spend, your geography and your conversion volume, not of ours, and it is worth knowing before you fund the campaign rather than after it. So all four verbs are graded at planning time rather than at reporting time. Delivered comes back as available once delivery starts. Responded comes back as needing a brief, needing instrumentation, not requested, or collecting. Caused comes back as eligible or not, with the reason. Where the design cannot be powered, the plan offers response and reach evidence instead and says that is what it is offering — which is a considerably less pleasant conversation to have at the point of sale than at the point of reporting, and a considerably more useful one.

Your spend

Concentration matters more than total. A test market has to carry enough of the buy for the exposed series to move at all against its own noise.

Your geography

A matched-market design needs markets to match. A national buy with no geographic dimension in its reporting leaves nothing to hold out and nothing to verify.

Your conversion volume

The days a design needs rise with the square of the baseline volatility and fall with the square of the lift you want to detect. Thin, noisy conversion series are the common reason a test cannot be powered.

Where the design cannot be powered, the plan says so and offers what it can evidence instead: delivery, response and modelled reach, each under its own label. That is a smaller claim, and it is the one the campaign can carry.

Ask the eligibility question before you commit the budget rather than after the flight, which is the only time the answer is still useful.

Talk it through

What buyers ask about measurement

Six questions, including the two nobody in this category volunteers: whether a causal result has ever been produced, and why a verb that has not been built is on the page at all.

Why publish a verb you have not built?

Because the alternative is to describe a three-verb model and let the fourth appear later as a launch. The four verbs are the measurement model; brand lift is the one that answers whether anything changed in people's heads, and leaving it out would make the model look complete when it is not. The API returns "remembered" with the status placeholder and a note that the study has not been commissioned, and the Cint adapter is non-transacting pending a contract. That is what this page says too.

What does it mean that the geo-lift engine refuses tests?

That if your spend, geography and conversion volume cannot detect a lift of a plausible size at 80% power, the engine returns an underpowered verdict and the number of days the design would need — and no lift figure. Most incrementality tools will happily return a number in that situation. The number is noise, and the harm is that it is indistinguishable from a real one once it is in a slide.

How long do you keep mobile advertising ids?

At most 90 days. After that the hashed id is persistently stripped from the exposure event, which re-grades the event to a coarser key rather than deleting the record. The rule lives in the key-translation service, which is the only module we have that performs identity joins at all.

Can you attribute a conversion to a billboard?

You can see a sequence, and that is a different claim. All seven channels write into one ledger, so the mechanism can show that a postcode sector exposed to a DOOH panel was later exposed to a CTV ad and later produced a conversion. Be exact about what feeds it: the ledger is fed only by the deterministic sandbox today, there is no inbound ingest route for viewshed panels, CTV logs or IP logs, and sandbox rows are labelled modelled rather than observed. So it is a working mechanism awaiting live input rather than a sequence anyone has yet seen in real delivery. Turning a sequence into a causal statement takes a powered geo-lift test in any case, and out-of-home is a good fit for one because the buy is bought by place.

Have you produced a causal result for anyone yet?

No. Nothing has run live, and a geo-lift is graded causal only when every geo outcome is observed, every delivery row maps to a live line and control exposure is verified — conditions sandbox data cannot meet by construction. Any lift figure you see in a demonstration is injected sandbox data labelled as a modelled demonstration. The engine, the power gate and the refusal are real and tested; the result is not yet evidence about a real campaign, and we do not print one as if it were.

Do you use marketing-mix modelling?

There is a modelled layer, and it is labelled modelled wherever it appears — deduplicated reach and frequency being the main case. What the product will not do is present a model output as evidence of causation. Modelled estimates and powered causal results are two of the four provenance labels precisely so that they cannot be read as the same thing.