M018 - Measuring AI When AI Changes the System
When the optimiser becomes a participant in the experiment, only Money In, Money Out survives.
The Asymmetry That Broke Modern Measurement
Measurement used to assume something quiet but load-bearing. The system you were measuring stayed put long enough for the measurement to mean something. Yesterday’s numbers explained tomorrow’s because the underlying relationships held. The whole infrastructure of modern marketing measurement sits on top of this assumption.
AI broke it.
The break is not the one most teams notice first. The platforms are not hiding the data. The signals brands send into the platform are stable, or in many cases improving. Server-side tagging is more mature than it has been in years. Consent infrastructure is in better shape. The platform’s view of the customer keeps getting better.
What changed is what the brand is able to read on the way back out. The outcomes the platform reports are now the outputs of a system rewriting its own behaviour faster than the measurement frame was built to handle. The signal in is consistent. The outcome read is fundamentally different.
M016 surfaced platform signals as the measurement everyone uses but will not admit. M017 looked at how to read PMax and ASC when the internal mechanics are hidden. M018 takes the next step. The system itself is no longer stable. And the measurement frames built around a stable system now produce confident answers to the wrong question.

The Stationary Assumption That Quietly Underwrites Everything
Walk through any standard marketing measurement conversation and the same assumptions show up.
Spend £1m on TV and generate roughly X sales. Increase paid search by 20% and get Y more leads. Audience A converts better than Audience B. The exact numbers move, but the relationships between the inputs and the outputs are assumed to hold.
This is the stationary assumption. MMM depends on it. Attribution depends on it. Forecasting models depend on it. Even many incrementality frameworks depend on it. The whole stack rests on the idea that what worked last year is informative about what is happening this year.
Most marketers do not notice how much of their measurement infrastructure sits on top of this single assumption. The signal hierarchy, the planning rhythm, the way trading reviews are run, the way budgets are defended at the board, all of it presumes a stable underlying system.
The assumption held when humans were the optimisers and the system moved slowly. The moment the optimiser became the platform, that survival window started closing.
What Changes When the Optimiser Joins the Experiment
The structural reason measurement breaks under AI is not about hidden data. It is about a feedback loop.
AI learns. When it learns, it changes its behaviour. When it changes its behaviour, it changes the environment the next decision is made against. When the environment changes, the historical data becomes less representative of the future. The next round of optimisation starts from there.
This loop is invisible from inside the planning meeting. Performance dips and the analyst goes looking for the campaign change that explains it. There isn’t one. The campaign is the same. The optimiser inside the campaign has moved, and that move shifted the conditions the previous quarter’s number was generated under.
The model is now a participant in the experiment, not merely an observer.
A useful way to picture this is GPS. Traditional marketing is driving the car. Measurement tells you whether you are still on the road. AI marketing is the car learning how to drive while you are driving it. Every decision changes the future behaviour of the car. The map itself is moving.
When the map is moving, last year’s road tells you almost nothing.
Three Worlds of Measurement
It helps to think about measurement as living in three different worlds.
Static Systems. The system changes slowly. Historical data remains useful. Classic MMM works well. This is the world most measurement methodology was originally designed for.
Dynamic Systems. The system changes regularly. Models require frequent recalibration. Most modern digital marketing lives here. Quarterly MMM refreshes, monthly attribution re-fits, ongoing incrementality programmes.
Adaptive Systems. The optimiser is learning continuously. The act of optimisation changes the future behaviour of the system itself. Historical relationships decay rapidly. This is the AI era.
The mistake teams make is to treat Adaptive Systems as a faster version of Dynamic Systems where the same tools work harder. They do not. A Dynamic System rewards more frequent recalibration. An Adaptive System needs a different discipline, because the thing being recalibrated keeps moving the calibration target.

The Cookieless Test That Almost Got Killed by Its Own MMM
The cleanest example sits inside a large UK digital programme.
A cookieless display intervention was tested using an incrementality design. The experiment ran clean. The intervention delivered real, measurable lift on incremental sales.
A quarterly MMM was then commissioned. The MMM was built on a historical window mixing two regimes. The earlier part of the window had no cookieless campaigns running. The later part had the cookieless campaigns active. The MMM smoothed across the regimes and returned a softer display performance read.
The recommendation it produced: reduce investment behind display.
Had the team followed the MMM as written, they would have cut spend behind something an experiment had already proven was working. The intervention was real. The model did not see it because the historical baseline was averaging two different systems together.
This is non-stationarity defeating the historical model in real time. It also reinforces the point M015 made about triangulation. The experiment said yes. The MMM said no. Without the discipline to triangulate, the wrong reallocation gets defended at the board.
Why Each of the Three Methods Bends Differently
Each method breaks under non-stationarity in a different way. The differences matter, because triangulating well means knowing which lens to trust under which conditions.
MMM assumes historical relationships are informative. When AI improves targeting, changes auction participation, finds new audiences or alters creative delivery, last year’s response curve no longer applies. The growth curve itself moves. MMM is the most affected because it is the most backward-looking.
Attribution assumes paths remain comparable. AI changes who sees ads, when they see them, the creative sequencing and the bid strategies. The path today does not resemble the path from three months ago. Attribution drifts continuously because the journey itself is no longer reproducible.
Platform signals are the trap most teams walk into. The metric labels stay consistent. Visits look like visits, ROAS looks like ROAS. The system underneath has changed, so period-over-period comparison is comparing two different systems wearing the same metric label. Performance changes get blamed on the wrong cause. Strategy is held responsible for something AI behaviour shifts produced.
Incrementality survives better. It asks what happened because of the intervention. It does not care as much about the underlying mechanics. It measures the difference between test and control at the moment of the test. That moment is stable even when the system is not. This is why experimentation becomes more valuable as AI adoption deepens, not less.
Money In, Money Out: The Frame That Re-anchors
M005 introduced Money In, Money Out as the only metric that ultimately matters. M018 is where MIMO does the work it was built for.
MIMO is built on measuring incrementality at the current point in time, then comparing investment against what is being delivered. It looks at the current point, not at the historical baseline. It is not impacted by lag behaviour, because each cycle re-anchors.
Where incrementality is available, MIMO uses the experiment as the baseline. Where it is not available, MIMO rebases using triangulation across the three methods. The re-anchoring is the whole point. MIMO does not depend on the historical relationship holding, because it keeps rebuilding the relationship from the present forward.
MIMO is not a fourth method. It sits above the three as a governance layer. The methods produce outputs. MIMO is the discipline translating those outputs into the return-on-investment language the business uses. When financial language is brought to marketing measurement, the conversation moves from individual measurement solutions to a MIMO framework.
Operationally, MIMO shows up in three places. In weekly trading, MIMO + triangulation as the read on what is happening. In quarter, when a new tactic or AI solution is introduced, MIMO + incrementality on whether the new thing works and at what level. In planning, MIMO + triangulation of existing information for forward allocation. Repeated for every product, every quarter, across all three surfaces.
The repetition is what gives MIMO its tolerance for non-stationarity. It is not solving the problem of a moving system. It is matching the cadence of the system’s movement.
When Prospecting and Retargeting Stop Being Separate
Two examples make the abstract concrete. The first is the collapse of prospecting and retargeting as separate tactics.
The old measurement was clear. Prospecting was measured on impression delivery and CPM. Retargeting was measured on pool size and CPA. Two creative sets were maintained. Budgets were separated. Performance was read separately. Everyone understood retargeting would have a better CPA than prospecting, and the optimisation question was the right mix between the two.
ASC collapsed the separation. Running retargeting alongside ASC would have doubled down on the same people, undermining both. The two strategies merged. Performance was measured at the joint campaign level. Creative aligned. Budgets combined.
The lingering trap shows up in the measurement infrastructure. MMM still studies prospecting and retargeting separately, assigning different incrementality levels to each. Weekly trading reports still split MIMO across them. Two lines move differently on the report, the team investigates, and finds nothing real to explain the variance. The variance is a measurement artefact, not a strategy signal. The two lines are responding to the same underlying cause, and the infrastructure has not caught up to the strategy.
What to retire and what to rebaseline. Retargeting efficiency stops being a measurable thing once ASC is running, so the metric should be retired. CAC and ROAS survive but only after a full re-baseline against the new combined campaign reality. AI adoption is a measurement reset event, not an incremental change. Treating it as incremental is what produces the phantom variance.
The deeper point lands here. When the tactic level breaks, planning moves up a level. The MIMO read is at total marketing. The strategic levers (product, offer, pricing, messaging, ecosystem) become the real control surface. Marketing consolidates back to its original levers. The AI takes the tactic. The marketer returns to strategy.

ASC, Reach and the Curves That Don’t Converge
The second example is from the growth curves work inside a large UK digital programme.
ASC, BAU and Reach were tested at multiple budget levels. The three behaved differently from one another. ASC kept giving returns as budget scaled. BAU saturated. Reach showed minimal diminishing returns. Three diverging curves, all inside the same programme, all measured the same way.
The MIMO move was to stop trying to reconcile the curves at tactic level. The question became simpler. Were both reach and conversion being covered. BAU was doing neither and was saturating. The money moved.
The reading underneath. ASC is a conversion engine. It finds better-converting prospects when the right signals are passed through. Reach is the net new prospects engine. It covers the audience the conversion engine has not yet seen. This maps cleanly back to D003: the only two strategies that matter are reach and conversion.
Audience overlap is where most planners get stuck and the wrong answer is to try to disentangle it. The controls held the system, not the audience. The control segment of ASC budget level 2 still received ads from other tactics, and so did the test segment. What was measured was the marginal contribution of that tactic, with overlap cancelling across test and control.
The honest caveat. AI optimises toward more likely converters, who are more likely to see ads on other platforms too. The lift measured includes that bias. This is not a flaw in the method. It is the structural reality of measuring tactics inside an ecosystem of overlapping AI optimisers.
The CFO conversation under MIMO works differently from the old one. The growth curves become the artefact. Where the programme is not at diminishing returns, the recommendation is to keep investing. Where it is, the recommendation is to scale back. Point estimates at each budget level give the bounds on confidence. The conversation is familiar enough to land (it echoes the old prospecting/retargeting trade-off) but the decision basis has moved from tactic-level CAC to curve-level marginal return.

PMax and the Generalisation Across Ecosystems
PMax pushes the same logic into Google’s ecosystem.
PMax is best treated as Google’s conversion equivalent of ASC. It combines Search, YouTube, Display, Gmail and Maps to find the best use of media investment. Demand Gen is the Reach equivalent.
The cross-surface delivery inside PMax is the strongest single picture of strategy and ecosystem collapsing into one AI-managed surface. The funnel as a planning construct disappears. The surface and channel decisions move inside the optimiser.
The measurement architecture transfers directly. The same 2-budget-level design that worked for ASC works for PMax. The frame generalises from one ecosystem to two, and the implication is that it generalises further as more platforms move to the same model.
The New Role of the Marketer: Periodic Intervener
A subtle asymmetry sits underneath all of this. The brand-side measurement models become less reliable under non-stationarity. The platform-side optimisation models become more reliable as they learn. The system is getting better at running itself while the system around it gets worse at reading itself.
This shifts the marketer’s role. The work becomes a small but high-leverage loop. Test if a new tactic works. If it works, estimate at what level of investment. Realign the media plans and marketing strategies. Let it run. Re-test periodically to check the assumptions or levels have not moved too much.
The marketer is no longer the continuous optimiser. The marketer is the periodic intervener and recalibrator.
This is where commercial judgement does work no method replaces. Deciding which tactics to test. Deciding when to re-test. Deciding what counts as enough drift to trigger a recalibration. Deciding which strategic lever to pull when the tactical lever is gone. Models do not make those calls. Marketers do.
Signal Stability vs System Stability
One distinction worth planting before closing. Stability of the signal is not the same as stability of the system.
Platform signals look stable in form. The metric labels and definitions hold across periods. The system the signals describe has shifted underneath.
Experimental signal looks unstable from one test to the next. The estimates wobble inside the bounds of statistical noise. The system being measured is stable, the noise is the limit of the method.
These are different problems. They demand different responses. Mistaking signal noise for system change leads to over-recalibration. Mistaking system change for signal noise leads to confident decisions against the wrong baseline. This distinction will return in later chapters. For now, the point is to notice that what is moving (the signal or the system) determines what the right response is.

The Discipline That Survives a Moving System
The future challenge is not measuring AI. The future challenge is measuring in systems where the act of measurement and optimisation continuously changes the thing being measured. This is a different problem from the one most measurement teams were built to solve.
The shift moves measurement away from prediction, reporting and attribution. It moves toward experimentation, signal monitoring and adaptive decision systems. The methods do not disappear. The frame that organises them does.
For the senior marketer reading this and deciding what to do tomorrow morning:
STOP relying on old benchmarks and strategies as the platforms and marketing ecosystems evolve.
START redesigning teams and their objectives to align with this changing behaviour, becoming intervener and recalibrator.
CONTINUE testing new tactics and investing in marketing analytics to build new strategies.
The growth curve is where the next chapter picks up. Before you understand headroom, you need to understand the curve itself might be moving. That is where M019 begins.
If you are trying to translate these mental models into operating systems that run day to day, this is the problem KaiSignals works on.





