Skip to main content
A measurement lead studying a wall of connected agent decisions in a dark room.

MMM Is Back. Agents Broke It.

Marketing is retreating to mix modeling at the exact moment autonomous spend makes its assumptions unstable. The fix is not to declare MMM dead. It is to operate it differently.

By Dellon S.12 min read

What MMM actually measures

Marketing mix modeling estimates how channels contribute to outcomes from aggregate spend, exposure, and business data. It does not need to follow every person or assign credit to one click.

MMM is having a renaissance because attribution lost signal. Google open-sourced Meridian, Meta maintains Robyn, and demand has climbed.

Now autonomous spend breaks the fallback underneath: the data window is too slow, channel meanings churn, platform metrics become targets, and calibration becomes circular.

MMM is the aggregate method that estimates channel contribution from spend and outcome data. The biggest signal is not a new model. It is a large advertiser admitting the old measurement frame can no longer steer the next dollar.

The $2B admission

In 2026, Adweek reported that Hershey is partnering with analytics platforms Mutinex and Tracer to modernize and automate its marketing mix modeling around roughly $2 billion in spend. The named case matters because it is not a startup demo. It is a mature advertiser acknowledging that its measurement system needs a new operating model.

Read the decision in two halves. First, a sophisticated MMM advertiser is saying the instrument no longer gives leadership enough confidence about where the next dollar should go. Second, the proposed cure is more autonomy, agents rebuilding the instrument that is supposed to supervise autonomous spend.

That may be a smart response to the speed mismatch. It is not automatically a smart response to opacity. A model that refits faster can still learn from platform metrics, unstable channel definitions, and decisions nobody logged. The real question is whether the new system adds independent evidence or only accelerates the loop.

Hershey is a public case, not proof that every agentic MMM program will fail. It is useful because it exposes the tradeoff early: if agents move the budget, measurement has to move faster, but it also has to remain outside the system's self-reported score.

A measurement team reviewing abstract model cards and colored signals at a dark table.
The measurement review is a decision room, not a dashboard tour. Someone still has to ask which assumption moved.

The renaissance is walking into the blind spot

Attribution's decline is not mysterious. Apple's App Tracking Transparency framework changed user-level signal in 2021, platforms walled off more data, and AI-mediated journeys made click-path logic less reliable. MMM became attractive again because it works with aggregate inputs and does not require a perfect user identity graph.

Google's Meridian framework and Meta's open-source Robyn have helped move MMM from a legacy CPG specialty toward a default enterprise fallback. Sellforte's Google Trends analysis describes demand accelerating since 2025, while EMARKETER's 2026 coverage tracks the same shift in analyst language.

The timing is the problem. MMM was rebuilt for a world where humans move budgets slowly and channel definitions remain stable long enough for coefficients to mean something. The agentic layer is a world of bidding agents, budget reallocators, creative optimizers, and cross-channel orchestrators making decisions hourly or daily. Marketing is retreating to an older instrument exactly as machines change the conditions that made the instrument legible.

The assumption to challenge is not simply “MMM needs more data.” It is that historical data still represents the policy a team is running now. A coefficient can be stable across a long history and still become misleading after an agent changes the bid rule, inventory mix, audience definition, or budget constraint. The old window contains signal about the market, but it also contains signal about decisions the current system no longer makes.

That changes what a measurement review should ask. Instead of treating a model refresh as a neutral update, separate three events: the market moved, the spend policy moved, or the measurement instrument moved. A model that cannot show which event explains a coefficient change is not necessarily wrong. It is simply too compressed to govern an autonomous system. The next version needs enough history to estimate durable effects and enough recent context to show where the policy has broken continuity.

This is also why the comeback should not be framed as a choice between attribution and MMM. The useful question is which claims each instrument can support. Aggregate modeling can estimate the contribution of a changing mix. Holdouts can test whether an intervention created incremental demand. Decision logs can show what the machine actually changed. The operating model has to join those claims instead of asking one technique to stand in for all three.

Editorial timeline comparing a 12 to 24 month MMM learning window with daily and hourly agent decisions, anchored by incrementality holdouts.
A model can be statistically careful and still describe a policy that no longer exists. Match the refresh cadence to the decision cadence, then anchor it to an external test.

What breaks when agents touch the mix

None of these mechanics says MMM's math is broken. They say the assumptions underneath the math are being managed by systems with a different objective.

01

The speed mismatch

MMM infers effectiveness from months, ideally years, of spend and outcome history. An agent can re-weight channels daily, so the channel estimate becomes an average over hundreds of policies that no longer exist. The model is a careful portrait of a budget that has already left the building.

02

Channel semantics collapse

MMM regressors use buckets such as search, social, retail media, and video. Agents dissolve those buckets by moving spend between placements, formats, inventory types, and platforms inside one control loop. The recommendation still says “shift 10 percent to social” after the category has been hollowed out.

03

Goodhart's law, automated

Give an agent a metric and it optimizes the metric, not the intent. Tell it to maximize CTR and it can find cheap inventory. Tell it to maximize platform ROAS and it can arbitrage the attribution window. The resulting correlation enters the model as if it were customer value.

04

Metric circularity

The spend agents optimize metrics computed by the platforms they buy from. MMM teams then calibrate against the same platform readings. Seller-graded homework, twice over. Without an external anchor, the stack can converge confidently on the platform's version of reality.

The contaminated outcome problem is not hypothetical. The related synthetic-data attribution analysis describes how machine-generated activity can look like demand. MMM inherits that contamination at the aggregate level, where it is harder to spot.

What policy generated this data? Record the agent version, objective, constraints, and major reallocations before interpreting the coefficient.

Which part of the result survives an incrementality test? A model can rank channels usefully while still overstating the lift that a platform claims.

Where do the model, experiment, and agent record disagree? That disagreement is a diagnostic signal, not a formatting problem to hide in an appendix.

The four mechanics compound. A daily reallocation changes channel semantics; the new semantic bucket is optimized against a platform metric; that metric is fed back into the model as evidence; and the model then approves another reallocation. Each step can look locally reasonable. The failure appears at the loop level, where the system is no longer measuring demand independently of the behavior it is rewarding.

Two black boxes stacked

The promise of agentic MMM is real. If spend moves at machine speed, a model that refits and recommends on a faster cycle is more useful than a quarterly report. The recursion problem is just as real: the instrument that audits agent spending is itself an agent, trained on the same platform metrics and capable of its own drift.

Diagram showing a spend agent feeding platform metrics into an MMM agent, with incrementality tests as an external anchor.

An agentic MMM that refits weekly genuinely addresses mechanic one. It deepens mechanics three and four unless it is deliberately caged. The measurement agent can learn from the same metrics the spend agent was rewarded for improving, then return a more confident version of the same circular story.

That is why the decision-ledger approach matters. Keep the model output, the agent decision, the policy version, the constraints, and the external check in the same record. If a coefficient moves, the team should be able to ask whether the market changed, the policy changed, the data changed, or the model changed.

Run MMM differently

Five changes separate MMM that survives agents from MMM that decorates them. They are not a new dashboard package. They are controls around the model's assumptions.

01

Model agent-managed spend separately

Stop pretending agent-allocated budget is a normal channel. Split regressors by control regime, human-planned versus agent-optimized, and separate platforms where the behavior differs. This lets the model estimate what delegation itself does instead of letting channel coefficients absorb policy churn.

02

Match refresh cadence to reallocation cadence

If agents re-weight weekly, quarterly refits are archaeology. Faster cycles are useful only when each refit is re-validated against holdout truth. A new model run is not evidence by itself.

03

Anchor calibration to incrementality

Use geo holdouts, blackout tests, and designed experiments as the data agents cannot Goodhart because they measure absence. Run them on a standing calendar and calibrate both the MMM and spend-agent objectives against experimental lift.

04

Join decision logs to model inputs

When the model sees a coefficient move, the decision log should show whether the policy, audience, bid rule, or tool changed. Without that join, every update arrives as an unattributable surprise.

05

Keep one human reconciliation ritual

Once a month, one page: what the MMM says, what the experiments say, what the agents did, where they disagree, and which number the next dollar follows. The ritual is deliberately boring. It is also where the boxes are forced to testify against each other.

The first change is the most important because it gives the rest a place to attach. If “agent-managed” is not represented in the data, the model cannot distinguish a channel effect from a policy effect. That distinction matters when a platform changes its bidding behavior without changing the campaign name. A clean campaign taxonomy can hide a completely different control regime.

The second and third changes create a clock and a referee. The clock says when a historical estimate becomes stale. The referee is the experiment that can disagree with both platform reporting and model output. Neither has to run at the same frequency as an individual bid. They do have to be scheduled before the next budget decision, not after a quarter has already been explained away.

The last two changes make the system inspectable. A decision log preserves the path from objective to action, while reconciliation preserves the disagreement between what the system did and what the evidence supports. Together they turn MMM from a periodic score into a control surface: a place where leadership can change the objective, the constraints, or the permission boundary before the next automated loop compounds the error.

Weekly

Review policy changes, not just spend totals. Name the new objective, the constraint that moved, the channels or placements affected, and the time window in which the model should stop treating old behavior as current.

Per test

Pre-register the holdout, the decision it will inform, and the threshold that changes the agent's behavior. If the experiment is only reviewed after the budget has moved, it is evidence of history rather than a control on the next decision.

Monthly

Reconcile the model's contribution estimates with observed lift and the agent's decision trail. Preserve the disagreements. A clean narrative that cannot explain its exceptions is weaker than a messy record that shows where uncertainty lives.

The point is not to slow every decision down for a committee. It is to keep decision speed and evidence speed separate. Agents can continue to bid, allocate, and test at machine cadence while the measurement system maintains a slower layer of independent checks. That separation gives leadership a place to challenge the objective without pretending the operating system is static.

The measurement record should connect to the wider control system. The holdout discipline for agentic ROI and the audit-trail approach are not separate governance projects. They are the external memory that keeps the mix model from becoming a polished guess.

That record should answer a simple handoff question: if tomorrow's operator inherits the budget, can they see why the system trusted the current mix? If the answer requires reconstructing a platform dashboard, an agent prompt, and a quarterly model file separately, the organization has evidence but not control. Joining those artifacts is the practical definition of measurement readiness in an agentic stack.

The handoff also needs a clear stop rule. When the model and holdout disagree beyond the agreed tolerance, pause the automated reallocation, preserve the inputs, and review the policy that produced them. That is a narrower intervention than shutting down every agent, and it gives the organization a repeatable response to uncertainty instead of a late explanation.

Two people comparing physical measurement cards and marked-up evidence at a dark review table.
A holdout is useful because it creates an absence the platform cannot award itself. Keep that external check close to the model, not in a separate quarterly appendix.

FAQs

Why is marketing mix modeling making a comeback?+

Attribution lost signal as Apple App Tracking Transparency restricted user-level tracking, platforms walled off data, and AI-mediated journeys broke click-path assumptions. MMM is aggregate and privacy-durable, so Google Meridian, Meta Robyn, and renewed industry demand have pushed it back into the center of the measurement conversation.

How does agentic AI break MMM?+

Agents reallocate spend faster than MMM training windows assume, dissolve stable channel definitions, optimize the platform metrics they are given, and create circularity when the measurement model is calibrated against those same platform metrics. The model can still be mathematically sound while its operating assumptions have changed.

What is Hershey doing with agentic AI and MMM?+

Per Adweek, Hershey is partnering with Mutinex and Tracer to rebuild marketing mix modeling around agentic AI for roughly $2 billion in marketing spend. The case makes the tension visible: agents are being used to measure a spend environment that agents helped make harder to measure.

Is agentic MMM a good idea?+

It can address the speed mismatch, but it adds another layer of inference. The measurement agent needs independent calibration, decision logs, sampled human review, and incrementality experiments. A faster refit is not the same as a more truthful result.

How should teams run MMM alongside AI spend agents?+

Model agent-managed spend as its own regressor family, match refresh cadence to reallocation cadence, anchor calibration to standing incrementality experiments, join agent decision logs to model inputs, and hold a recurring human reconciliation of the model, experiments, and agent behavior.

Abstract measurement signals moving through a dark field of intersecting paths.

MMM earned its comeback because it is durable when identity and platform data fail.

The agentic era does not kill the model. It kills the old assumptions around it.