Skip to content

Glossary Incrementality test

What is an incrementality test

Definition

An incrementality test is a controlled experiment that compares a group exposed to advertising with a comparable group that is not exposed, in order to estimate how many conversions the ads actually caused and how many would have happened anyway.

On this page 5
  1. What incrementality means
  2. How an incrementality test is set up
  3. Why it matters
  4. Good practice
  5. Common mistakes
In brief

An experiment with a control group that separates the conversions caused by advertising from those that would have occurred without it.

What incrementality means

Attribution distributes conversions across the touchpoints a user hit before buying. It answers a descriptive question: which path that sale took. Incrementality answers a different and considerably more uncomfortable one: would that sale have happened without the ad?

A customer who already knew the brand, who was going to buy anyway and who clicked an ad along the way shows up in the reports as a paid conversion. The budget takes credit for a result that existed before the money was spent. The incremental effect is only the part of the result that exists because the investment was made.

Measuring that part requires a comparison with a scenario that cannot be observed: what would have happened without the campaign. Since that scenario is not in the data, it has to be constructed. A comparable group is set aside and the ads are withheld from it, and its behaviour serves as the reference. The difference between the two groups is the incremental effect.

Hence the name. What is measured is not how much activity happens around the ad, but how much activity the ad adds.

How an incrementality test is set up

Every design needs two things: a group that sees the advertising and a comparable one that does not, plus a random assignment that decides who ends up on which side.

Holding back part of the audience is the most direct design. The platform randomly sets aside a share of eligible users and hides the campaign's ads from them. Custom experiments in Google Ads work on that logic: the cookie-based split assigns each user to the original campaign or to the test campaign and guarantees that they only see one of the two.

The geographic experiment replaces the user with the territory. Regions with similar sales histories are paired up, a draw decides which of each pair keeps its investment and which one reduces or switches it off, and the total sales of both blocks are compared. This design depends on neither cookies nor identifiers, so it survives the loss of user data, and it is the standard choice for channels where the click cannot be tracked.

Switching off by period alternates weeks with investment and weeks without it in the same territory. It is the cheapest design and also the most fragile, because any seasonality or competitor move gets confused with the effect you are trying to measure.

What comes out of the test is a difference between groups: additional conversions, additional revenue or, divided by the investment, an incremental return. That figure always comes with a confidence interval, and the interval matters as much as the central value.

Why it matters

Almost every budget decision is made on attribution figures, and attribution does not distinguish between causing a sale and witnessing one. A channel can look excellent simply because it appears in the last stretch of the journey, once purchase intent was already formed.

Brand search campaigns are the case where that gap becomes most visible. Someone searching for the company name has already decided where they are going. The ad places itself in front of an organic result that would have received the same click, and the report records a paid conversion with an enviable acquisition cost. An experiment that switches those campaigns off across a group of territories shows how much of that demand merely moves from the free channel to the paid one. There are published experiments in which brand keywords showed no measurable short-term benefit, although the outcome depends on the sector, on brand awareness and on whether a competitor is bidding on that name.

That is why the result of an incrementality test almost always comes out below the one any attribution model offers. It is not that the experiment punishes the channel. Attribution never promised to measure causality.

Good practice

  • Decide before you start what difference would be relevant to the business, and calculate how long the test has to run to be able to detect it.
  • Pair territories or groups by sales history, not by population size.
  • Leave the campaign alone while the experiment runs: no changes to bids, creatives or budgets in either arm.
  • Measure the result in the business source, meaning the ERP or the CRM, and not in the platform being evaluated.
  • Report the confidence interval alongside the result and accept that an interval crossing zero is a legitimate answer.
  • Repeat the test from time to time, because a channel's incrementality changes with brand awareness and with competitive pressure.

Common mistakes

  • Switching a campaign off for a week, seeing that sales do not drop and treating the matter as settled. Without a control group there is nothing to compare against.
  • Stopping the test as soon as the result is pleasing. Checking daily and stopping on the first favourable day turns noise into a conclusion.
  • Choosing a window that is too short for long purchase cycles, so the delayed effect falls outside the measurement.
  • Contaminating the control group with remarketing, campaigns from another channel or emails that do reach those users.
  • Transferring the result to another country, another season or another investment level without measuring again.
Manuel Riveiro Rodriguez CEO & Digital Strategist

A technical audit covers this and everything else in one pass.

Request an audit

Frequently asked

How does an incrementality test differ from an A/B test?

An A/B test compares two versions of an ad or a page and says which one performs better. An incrementality test compares advertising against the absence of advertising and says whether the channel contributes anything at all. The first answers which variant to choose; the second, whether to invest.

How long does an incrementality test need to run?

It depends on conversion volume and on the size of the effect you want to detect, not on the calendar. As a practical guideline it should cover at least one full purchase cycle and several weeks of stable data. A power calculation up front avoids running a test that can never conclude anything.

Why is the result lower than in Google Ads or Analytics?

Because they measure different things. The platform counts conversions that passed through the ad; the experiment counts those that would not have existed without it. The second figure is always equal or lower. The difference is not a tracking failure, it is the demand that was already there.

Can incrementality be measured without cookies?

Yes. Geographic designs assign the condition to entire regions and compare aggregated sales, so they need to identify no user and do not depend on cookie consent. That is why this format has gained weight since individual tracking became less reliable.

Does an incrementality test replace attribution?

No. Attribution remains useful for observing journeys and optimising day to day, because it is always available and highly detailed. The experiment provides the reference that corrects that reading from time to time. The sensible approach is to steer with attribution and calibrate with the experiment.

Sources

  1. Documentation of custom experiments in Google Ads: the cookie-based split randomly assigns each user to the original campaign or to the test campaign and guarantees that they only see one of the two.
  2. Google paper describing the geographic design: non-overlapping regions randomly assigned to a control or treatment condition, delivered through geo-targeted advertising.
  3. Open Google library for designing and analysing geographic experiments with matched regions and estimating the incremental return on ad spend.
  4. A series of large-scale field experiments at eBay on the effectiveness of paid search; in the extreme case of brand keywords, no measurable short-term benefit was observed.
  5. Meridian guidance on using past experiments to calibrate a model, with the caveats on time window, duration, channel mix, compared scenario and population.