M017 - BLACK BOX NOT BLIND BOX: MEASURING WITHOUT INTERNALS
Every major ad platform has moved to opaque AI. The opacity is not the problem. What marketers do with it is.
There is a phrase practitioners have been using for years to describe the growing opacity of digital advertising. Black box. It started with Google’s Quality Score. It migrated to programmatic auctions. It found a permanent home with machine learning campaigns.
Performance Max. Advantage+ Shopping Campaigns. TikTok Smart+. Demand Gen. Snap Automated Ads. Pinterest Performance+. Every major platform has either launched an AI-driven campaign type or made it the primary recommendation. The direction is not reversing. More data processing, more signal interpretation, more automated decision-making -- the platforms are getting better at what they do, and the trade-off is opacity.
What these systems share is a common structure. You provide the inputs -- conversion data, audience lists, creative assets, product feeds -- and the system decides. Which audiences to target. Which placements to buy. Which creative to serve. What to bid. You see the outputs. You do not see the decisions.
This article is not an argument against that structure. It is an argument against what happens when marketers respond to it by stopping measurement altogether.
A black box is a technical reality. These systems process signals at a scale no manual setup is able to replicate. The opacity is the price of that processing power. Criticising it is like criticising a search algorithm for not showing you the source code. It is also, for many brands, the source of genuine performance improvement.
A blind box is something different. It is not a property of the system -- it is a posture adopted by the marketer. It is what happens when, because you are not able to see the internals, you also stop measuring the outcomes independently. You accept the platform’s reported results as capital allocation truth. You stop asking whether those results reflect genuine incremental value for the business. You hand the measurement to the thing being measured.
The core question this article addresses is a practical one: how do you measure systems you are not able to see inside? Not using platform reports. From the outside, using your own data, against your own business outcomes. That is what money in, money out means in practice, and it remains constant even when the system generating the returns is opaque.
(For the MIMO foundations, see M005. For the incrementality principles behind this, see M007 and M014.)

The Transparency Trap
In the past year, platforms have added meaningfully to their reporting. Google’s Channel Performance report shows where Performance Max spend went. Meta’s Incremental Attribution tool offers a view of incremental impact. Asset-level metrics give creative performance data that did not previously exist. The industry has largely treated these additions as measurement progress.
They are not. They are transparency progress. And the two are different things.
Transparency features tell you what happened inside the platform. Which placements received budget. Which creative ran most. How spend was distributed. That is useful for understanding the system’s behaviour -- in the same way a bank statement tells you how your money was spent without telling you whether the spending was worthwhile. The platform report shows movement. It does not show value created.
The deeper problem is one of muscle memory. Incrementality measurement -- proper causal testing -- has long lead times. It requires planning, infrastructure, and patience. In the absence of it, practitioners have always defaulted to platform metrics as a directional shortcut. Over time, the shortcut becomes the route. The directionality of platform metrics gets confused with the directionality of true incrementality. The two get conflated. Platform numbers get given more significance than they were ever designed to carry.
This is not a new problem. What is new is that platform transparency features accelerate it. They make the proxy more visible, more granular, and more convincing. The more it looks like measurement, the less reason practitioners feel to do actual measurement. They were always meant to be a bridge between what is readily visible and what is truly measurable. The mistake is building on the bridge rather than crossing it.
What Happens When You Trust the Platform
Platform-reported performance has a structural bias. Every platform measures its own contribution using the signals available within its own ecosystem. It does not see conversions attributed elsewhere. It does not deduct for overlap. It is not able to account for what would have happened without the ad. The result is systematic over-crediting -- not by intent, but by technical constraint.
When marketers use these numbers for capital allocation, the over-crediting compounds. They over-estimate performance and over-invest in what appears strongest on platform. Cross-platform comparison breaks down because every platform runs the same bias in the same direction. The platform gaining the most interactions captures the most attributed sales and garners the most investment -- regardless of incremental contribution.
Then the AI learns.
The AI’s objective function is informed by the outcomes the marketer rewards. It learns to find more of what gets rewarded: high-intent, low-incremental conversions. People already in the purchase funnel. Existing customers near renewal. Brand searchers with established intent. These users are easy to find and cheap to attribute. The system converges on them. Mid-funnel investment gets deprioritised because it generates poor platform-reported returns.
What follows is gradual and largely invisible. Sales begin to erode. Market share declines. No single dashboard connects the erosion to the original decision. By the time the weekly trading reports make the problem visible, the mid-funnel has been quiet for months. Recovery is slow. The pipeline does not restart quickly. It needs rebuilding from multiple directions over an extended period.
This is not a rare failure mode. It is the natural end state of treating platform metrics as capital allocation truth. It is a self-fulfilling circle: over-crediting drives over-investment, which trains the AI to optimise for what gets credited, which starves the mid-funnel, which erodes the pipeline, which erodes returns -- whilst the dashboard continues to report efficiency.
A Three-Stage Approach to Measuring Without Internals
Measuring a black box does not require seeing inside it. It requires three questions, asked in sequence.
Stage 1: Can I execute?
The first question is whether the new campaign type is able to run your strategy at all. Is business-as-usual execution achievable through this system? This is a binary feasibility check. It precedes any performance judgment. You are not yet asking whether the system works. You are confirming you are able to use it without breaking what already runs.
Stage 2: Is it detrimental?
The second question is whether switching to this approach causes measurable harm to total business outcomes. This is still a binary question: directionally damaging, or not. The answer comes from weekly trading reports monitored over a sustained period -- not a single week.
Two things matter here. First, allow for stabilisation. Platform AI requires a learning period before performance settles -- typically two weeks, but variable depending on traffic and conversion volume. The high variance in early days is noise. The measurement clock starts only after the learning period has ended. Second, look for a sustained direction, not a single data point. Directional monitoring requires several weeks of post-learning data before a judgment is reliable.
Stage 3: How good is it, and at what investment level?
The third question is the most important and the hardest to answer. It requires moving from directional monitoring to quantified incrementality. How effective is this system, genuinely? At what level of investment does it remain profitable? Where are the diminishing returns?
This question is not answerable using previous knowledge of the platform. AI changes the performance dynamics of a campaign type. The saturation curves from your previous manual campaigns no longer apply. You need to rebuild that understanding through formal incrementality testing -- a Conversion Lift Study, a geo-based test, a pulse study. Each method carries its own limitations, and those limitations need to be understood in the specific context of AI-driven campaigns.

Why Geo-Lift Tests Break for AI Campaigns
The natural instrument for Stage 3 is a geo-lift test. Split the country into test and control geographies. Run the campaign in test, withhold it in control. Compare sales outcomes. It is the cleanest money in, money out instrument available.
But it breaks for AI campaigns. The reason is structural, and it is worth understanding clearly.
AI-driven campaign types -- Performance Max, ASC, Advantage+ -- optimise across the entire available audience. Their performance is a function of the full signal set they are given access to. The moment you split the country, you are no longer asking the AI to find the best outcome across all potential customers. You are asking it to find the best outcome within a geographic subset. Those are different problems with different answers.
The AI in your test geography is not performing at its true potential. It is performing at the potential of a constrained version of itself, working within a limited pool of users. The test measures a degraded version of the strategy and draws conclusions about the real one. For Stage 1 and Stage 2 questions, this distortion is acceptable -- you are looking for direction, not precision. For Stage 3, it is not. You need to know how the strategy performs at scale, not at subset.
This is why the All In on AI measurement at Sky pivoted from a geo-lift to a Causal Impact study. The question being answered was Stage 3: at the current investment level, with the current strategy, was consolidation to AI-only working? A geo-lift would have answered a different question -- how does a constrained version of the strategy perform in part of the country? The Causal Impact study measured the effect of the strategic intervention against a historical baseline, without constraining what the AI was doing.
The acknowledged trade-off was wider confidence intervals. The analytical work required to interpret those intervals was significant. But a wide confidence interval around the true effect is more useful than a precise measurement of the wrong thing.
The Sky Case Study: All In on AI
In 2024, Sky ran a consolidation that is directly relevant to this argument. The brief was significant: reduce from 13 advertising partners down to 2 (Google and Meta), replace manual campaign structures with AI-only campaign types, and assess whether the consolidated approach delivered the same commercial outcomes at the same investment level.
The setup was a direct implementation of the framework from D003. Meta Reach handled the reach strategy -- expanding the pool of people who had been exposed to the brand. Meta ASC handled conversion -- finding the next purchasing consumer within the audiences reach had developed. Google Performance Max and Brand Search handled intent capture and conversion across search and display. Four campaign types. Two strategies. No manual intervention in targeting or placement decisions.
Measuring this required a deliberate decision to ignore platform reports. The setup included mid-funnel Reach campaigns alongside conversion campaigns. Mid-funnel activity generates post-view conversions: a consumer sees the ad, does not immediately respond, and converts later through a different touchpoint. Platform attribution is structurally unable to capture this accurately. Using platform numbers as the measurement instrument would have systematically undervalued the Reach campaigns and distorted the picture of the whole strategy.
The Causal Impact study compared the consolidation period against a historical baseline, using pre-existing trends to isolate what the intervention had caused. The limitations were acknowledged from the start: confidence intervals are wider than a geo-lift, and the estimation of actual performance requires careful analytical treatment before a judgment is made. That treatment was done.
The result was 12% sales growth at net-neutral budget.
That number did not come from platform reporting. It came from a measurement instrument to which the platforms have no access -- the brand’s own sales data, reconciled against a causal model of what would have happened without the intervention. That is money in, money out. Applied from the outside of a black box.
The Mid-Funnel Trap
There is a specific failure pattern worth naming because it recurs across organisations running AI-driven campaigns.
Mid-funnel activity -- Reach campaigns, Demand Gen, upper-funnel formats -- consistently receives poor performance ratings in platform reports. The structural reason is post-view attribution. These formats drive consumers to awareness and consideration. The conversion happens later, through a different touchpoint, and the attribution goes elsewhere. The mid-funnel campaign did the work. The lower-funnel campaign gets the credit.
The platform report says: this activity is not working, consider switching it off.
The correct interpretation is: this activity is doing something the platform is not able to see.
When organisations follow the platform report and switch off mid-funnel activity, the commercial failure follows with a lag. Sales erode gradually. Market share declines slowly. No dashboard connects the trading decline to the campaign decision made weeks earlier. The link is invisible because the mechanism was invisible. By the time the erosion surfaces in weekly reporting, the pipeline has been draining quietly for months.
Recovery is disproportionately slow. The audiences built through reach, the intent developed through upper-funnel exposure -- these take time to rebuild. Months, typically, with support from multiple channels simultaneously. The organisation ends up spending significantly more to recover than it would have spent to sustain.
The Haus incrementality study, covering 640 tests across 18 months, found Meta’s ASC outperformed manual campaigns in 42% of cases. That finding is not a verdict on AI campaigns. It is a diagnostic prompt: under what conditions does platform AI outperform manual, and what determines the outcome? Signal quality, creative volume, investment level, purchase cycle length -- these interact with the AI’s ability to optimise. Without measuring correctly, you never understand which conditions you are operating in. You get an average result and no lever to change it.

What Correct Measurement Unlocks
Measuring AI systems correctly from the outside is not a technical exercise. It is the foundation of strategic agency.
A team with genuine understanding of the true incrementality of its campaign types is able to build a cross-platform strategy tied to business objectives rather than platform metrics. The platform AI becomes an execution layer for a strategy the marketer owns -- not the other way around. The cross-platform picture becomes coherent because the measurement is consistent: money in, money out, assessed independently of what any single platform reports.
The fundamental levers return. Right customer. Right time. Right place. Right ad. Right offer. These did not disappear when platform AI arrived. They got obscured by the opacity and, to be honest, by the complacency the opacity encouraged. The intermediate period of platform AI adoption created two types of organisations: those confused by what they could not see, and those comfortable not looking. Both drifted from the fundamentals.
Measuring correctly is what brings them back. The AI handles targeting and optimisation decisions at a scale no marketer is able to replicate. The marketer handles the strategic design decisions the AI is not equipped to make -- what the strategy should be, how to measure it, when to constrain the system and when to let it run. That is the right division of labour. It is also what the relationship between brand and platform was always supposed to look like, before measurement got outsourced to the platform itself.

CLOSING: Reclaim the Position
Platform AI will get more capable and more opaque in parallel. That is the direction. More signals processed, more decisions automated, less of the internal logic visible. The transparency features will improve too -- but transparency and measurement are not converging. The gap between what the platform reports and what the business experiences will remain a gap the brand has to close for itself.
The answer is not to resist the opacity. It is not to accept it on faith.
Implement money in, money out. Understand the exact levels of media effectiveness these black box systems are delivering. It does not matter that you cannot see how they target and optimise -- as long as they are making media more effective by finding the right customer at the right time in the right place, and you have a way of knowing that independently. Measure whether they are. Build marketing plans on what the measurement tells you. Refine strategies as the evidence accumulates.
There is a transition period here and it takes time. Causal studies require planning. Confidence intervals are wider than a direct platform read. The analytical work is real. But a wider confidence interval around the truth is more useful than a narrow read of a proxy. The organisations building this measurement foundation now will have something their peers will not: a strategy they are able to defend at board level because they know what it is delivering -- not just what the platform says it is delivering.
The black box is given. The blind box is a choice.
M018 looks at what happens to measurement when the AI system starts changing the environment it operates within -- and why the standard measurement models begin to break down when the system under study is also shaping the signals it generates.
If you are trying to translate these mental models into operating systems that run day to day, this is the problem KaiSignals works on.





