Your attribution model used to be finicky. Now it's hallucinating.
Last week, a CMO at a mid-market SaaS company told me her Google Ads AI said she was getting 3x the conversions she actually was. Her team didn't notice for six weeks. By the time they caught it, her marketing operations analyst had already rebalanced the entire budget based on what turned out to be phantom ROI.
This isn't a story about bad AI. It's a story about what happens when your marketing models get trained on data that doesn't exist.
The Collapse Loop
Here's what's happening at scale right now:
AI models need training data. There isn't enough real data anymore, so companies use AI to generate synthetic training data. That synthetic data is statistically plausible but causally hollow. It looks right but doesn't reflect the world. These models get trained on the synthetic data. They then get deployed into marketing platforms. Marketing platforms use these models to generate recommendations. Those recommendations become decisions. Decisions become data. And that new data cycles back into the next generation of training data.
You're feeding your measurement systems garbage collected from models trained on garbage.
The math breaks. Attribution coefficients drift. Incrementality tests become unreliable. You lose signal without knowing you lost it.
It's not a bug. It's a structural consequence of closing the loop too fast.

Why Attribution Breaks Specifically
Attribution models are coefficient-hungry. They're trying to estimate how much credit each touchpoint gets for a conversion. This works when your training data reflects actual causal relationships. User sees ad, user acts, conversion happens. The signal is real.
But synthetic data doesn't have causality. It has correlation. An AI system generating fake user journey data will create statistically coherent sequences that have never actually led to conversions. The model learns to detect patterns that don't predict real behavior. When you ask it to attribute real conversions, it attributes them to the phantom patterns instead.
The result: you're maximizing toward signals that don't exist.
Google's own models are doing this. So are CDP attribution engines, Conversion API lookalike audiences, and next-best-action personalization systems. Every one of them is quietly training on synthetic data now. The feedback loop is already three layers deep.
Where You're Seeing This (And Not Seeing It)
If you're using any of these, you've got synthetic data in your training pipeline right now:
Google Ads Performance Max. Google trains these models on historical campaign data, then fills gaps with synthetic sequences to improve performance under cold-start conditions. They're transparent about this in documentation almost nobody reads. When your budget is deployed against audiences that were trained partly on invented user journeys, you're bidding on predictions instead of patterns.
First-party CDP attribution. Customer data platforms are using generative models to fill missing data points in journeys. A user lands on your site, bounces, re-enters through email, converts. The CDP's inference engine might synthesize a missing touchpoint in the middle to complete the sequence. That synthetic touchpoint ends up in your attribution model. Now you're attributing credit to interactions that never happened.
Analytics model refresh cycles. Analytics platforms backfill missing data using ML models trained on synthetic data. Some of this is disclosed in your contract. Most of it isn't.
The common thread: every system is trying to solve the cold-start problem by generating synthetic training data. And every system thinks its synthetic data is good enough for downstream decisions. None of them account for compound error across the stack.
The Budget Bleed You Won't See Until October
Here's what typically happens:
July / August: Attribution looks normal. You're getting 30-day ROAS of 3.2x. Looks good.
September: Your incrementality test shows 1.8x actual ROAS. Something's wrong. You run diagnostics.
October: You realize your attribution model has been trained partly on synthetic data for the last 60 days. By now, 40-60% of your budget has been allocated based on phantom signals. You're not losing money directly. But you're leaving 20-30% of potential upside on the table. And if you're in a competitive category with thin margins, that's the difference between market leader and also-ran.
For a $10M annual ad budget, 20% is two million dollars.
Most CMOs don't audit their model training pipelines quarterly. They should.

What CMOs Actually Do (Tactical Moves)
"Be careful with synthetic data" isn't actionable. Here's what moves:
Segregate your attribution audit. Build a separate cohort of real, human-verified conversions (not inferred, not synthetic, not ML-backfilled). Train a control model on only real data. Compare it to your production model monthly. If divergence exceeds 15%, your production model is training on too much synthetic data. Call your vendor and ask what changed.
Check your cold-start thresholds. Most platforms have a setting for how much synthetic data they'll use during model training. Google Performance Max has a disclosure in the API docs. It's not in the UI. Ask your account team. If it's >20%, lower it or turn it off and accept slower scaling.
Test incrementality on new audiences. When you launch a new audience targeting cohort, run a 50/50 holdout test for 14 days before full deployment. Incrementality tests are expensive but they're the only thing that catches synthetic data model drift early. If actual ROAS on holdout is 30% lower than predicted ROAS, your model is hallucinating.
Audit your CDP refresh cycle. Ask your CDP vendor: "What percentage of our journey data is synthetically backfilled?" If it's >10%, ask them to turn it off or lower the threshold. Yes, your data will have gaps. That's better than your data having fabrications.
What's Coming
Regulatory bodies are starting to care about this. The FTC issued a notice in Q2 2026 flagging AI training data transparency as a consumer protection issue. They haven't come down hard yet. They will.
Model obsolescence is also coming. As more synthetic data cycles into training pipelines, model performance on real data degrades. This is called "model collapse" in research. At a certain threshold, models stop learning from new real data because they're overfitting to statistical artifacts of synthetic data. When that hits your ad platform models, scaling stops. Budget doesn't convert anymore, regardless of bids.
This is still a few months out for most platforms. But you should be thinking about it now.

The Measurement Crisis You Can Act On Right Now
The core issue is this: marketing measurement is being trained on data that doesn't exist, validated against benchmarks created from other synthetic data, and deployed against your real budget. The gap between prediction and reality is widening. It's not visible in your weekly dashboards because dashboards are built from the same degraded models.
You need to measure outside the system.
Real incrementality tests. Real cohort audits. Real data-only model training. Real quarterly divergence audits.
It's work. But it's the only thing that catches the feedback loop before it catches your budget.
Start this week. Your Q4 ROAS will thank you.
