Skip to main content
A marketing strategist comparing different AI outputs across three monitors.

The Jagged AI Problem

Better models can make marketing worse when every request goes through the same lane. The fix is not blind upgrades. It is model routing, workflow evals, and proof by task.

By Dellon S.May 13, 202611 min read

The upgrade trap

The trap starts with a reasonable sentence: this model is smarter, so everything should improve.

Then the campaign calendar gets wordier. The ad variants start explaining themselves. The lead response is technically correct but arrives too late. The customer support reply is polite, complete, and somehow less useful than the old version.

The team calls it a prompt problem because prompts feel fixable. Rewrite the instruction. Add examples. Threaten the model with a style guide. Sometimes that helps. But if the wrong work is going to the wrong model, prompt polish only hides the routing mistake.

Marketing work is not one task. It is a pile of different jobs wearing the same department name. A positioning memo needs judgment. A product page needs voice. A subject line needs restraint. A lead response needs speed. A citation-backed article needs source discipline. A cleanup task may not need a frontier model at all.

Campaign materials being reviewed after a model upgrade creates inconsistent marketing work.
The upgrade trap shows up in the work before it shows up in the dashboard.

What jagged means

Jagged does not mean bad. It means uneven. A model can be brilliant at one kind of work and strangely clumsy at another. The same system that sees a strategic contradiction in a brand architecture deck can over-explain a six-word ad line. The same model that writes a careful competitive analysis can turn a direct customer response into a padded paragraph.

That unevenness is easy to miss because most model discussions happen at the wrong altitude. Teams compare models as if they are replacing a single machine. Marketing teams do not use AI as a single machine. They use it as a nervous system across briefs, ads, emails, landing pages, reporting, customer response, search research, content refreshes, and sales enablement.

That is why generic model rankings are not enough. Anthropic's model selection guidance tells teams to weigh capability, speed, cost, and effort before choosing a model for an application. OpenAI's evals documentation points in the same direction: test model behavior against a dataset and criteria that match the job. For marketing leaders, the important lesson is simple. The eval belongs to the workflow, not to the model hype cycle.

This is where a lot of AI marketing programs quietly go sideways. The first rollout looks promising because the team tests impressive tasks. Strategy decks. Long-form drafts. Competitive teardown. Brainstorming. Then the same model gets dropped into high-volume production. Now the team is using expensive reasoning for cleanup, slow inference for time-sensitive replies, and a model with a distinctive writing fingerprint for copy that needs to sound like the brand.

If you want a useful mental model, stop asking, "Which AI model is best for marketing?" Ask, "Which marketing job is this, and what would failure look like?"

Build the routing layer

The routing layer is the part of the system that decides where work goes before anyone argues about prompts. It can be code, a workflow rule, an automation, or a human review step. The form matters less than the discipline.

Marketing jobLikely laneQuestion before routing
1Campaign strategy

Reasoning model

Can it weigh constraints, tradeoffs, positioning, and risk without flattening the point?

2Email subject lines

Fast general model

Can it create tight variants quickly without explaining the joke?

3Research synthesis

Citation-first workflow

Can every important claim be traced back to the source?

4Lead response

Low-latency path

Can it answer accurately before the buyer cools off?

5Cleanup and tagging

Small model or rule

Can a cheaper lane handle it without touching brand voice?

Bad route

"Send every marketing request to the most capable model and trust the prompt to handle the rest."

Better route

"Send each task to the lane that matches its speed, risk, evidence, cost, and voice requirements."

A strategist sorting physical task cards into different AI workflow lanes.
Routing becomes easier to manage when the work has visible lanes instead of one default model.
A folder being stopped before it enters the wrong physical routing lane.

What breaks first

Jagged results rarely arrive as a clean outage. They show up as small losses that feel unrelated until someone maps the work by task.

Delay tax

The response is smart, but it arrives after the moment has passed.

Voice drift

The model changes sentence shape, confidence, and rhythm until the brand sounds borrowed.

Cost creep

High-volume tasks quietly move through premium models because nobody owns the route.

False confidence

A polished answer hides weak evidence, stale assumptions, or an untested workflow.

Measure by workflow

Model benchmarks are useful, but your campaign does not care about leaderboard wins. It cares whether the subject line ships on time, sounds like you, and improves the outcome that section owns.

A good AI model audit asks a boring question over and over: what does usable mean here? For a landing page refresh, usable may mean a sharper angle and less editing time. For research, usable may mean source fidelity. For lead response, usable may mean low latency and no hallucinated promises. For paid social, usable may mean strong variation without losing the brand's restraint.

WorkflowPrimary testWrong signal
Email subject linesOpen rate and brand fitLongest reasoning trace
Campaign strategyDecision qualityFastest draft
Lead responseSpeed and accuracyMost polished paragraph
Research synthesisSource fidelityConfident tone
Content refreshUseful additions and freshnessMore words alone
A marketer comparing charts, costs, and timing across campaign workflows.
If the audit only asks which model is smartest, it misses the tradeoff that changed the result.

A practical routing audit

The easiest way to find the jagged edge is to follow one real piece of work from request to result.

Pick a workflow that happens often enough to matter. A weekly newsletter. Paid social variants. Demo request follow-up. A blog refresh. Do not begin with the model picker. Begin with the moment where someone on the team says, "AI should handle this."

Then write down what the work actually needs. Some tasks need strong judgment because a wrong answer changes positioning. Some tasks need a clean voice because every sentence touches the brand. Some tasks need fast response time because a buyer is waiting. Some tasks need hard source control because the page will be cited, reused, or summarized by AI search systems.

Once the job is named, test the current lane against the outcome. The best audit is small and slightly uncomfortable. Use ten real examples, including the ones the model previously missed. Ask the same workflow to run through the current path and a challenger path. Score both outputs on the thing the workflow exists to improve, not on general fluency.

Step 1

Trace the handoff

Who starts the request, what context is included, which model handles it, who reviews it, and where does the output go next?

Step 2

Score the miss

Was the output slow, expensive, off-brand, unsupported, too generic, too cautious, too long, or simply aimed at the wrong business result?

Step 3

Change one lane

Move one workflow to a better route, keep the old path available, and compare cost, speed, editing time, and usable output for a full cycle.

The point of the audit is not to embarrass the model. It is to remove mystery from the system. If a subject line route fails because the model keeps explaining itself, the fix may be a faster model, a shorter prompt, or a hard length rule. If a research route fails because citations are weak, the fix may be retrieval, stricter source acceptance, or a human editor before publication. Those are different problems, and they should not share the same solution.

This matters even more for GEO work. AI answer engines often compress a page into a few extracted claims. If your workflow uses a model that sounds confident but drops source context, the page may read well to a person and still become hard for machines to quote accurately. That is a routing issue. Citation-heavy work needs a lane that protects evidence, names entities clearly, and keeps claims close to support.

A good routing audit gives the team a shared language. Instead of arguing whether one model is "better," people can say the strategy lane improved, the subject line lane slowed down, the research lane needs stricter citations, and the cleanup lane is wasting money. That is the moment AI becomes operational instead of theatrical.

The operating model

1

Name the job

Do not start with the model name. Start with the workflow and the outcome it owns.

2

Pick the lane

Route by reasoning need, latency, cost, evidence standard, and voice sensitivity.

3

Write the eval

Create a small test set for each workflow. Use real briefs, real customer questions, and real failures.

4

Version the route

A model change should be treated like a product change, with owner, date, reason, and rollback path.

5

Review the misses

Look at bad outputs weekly. The mistake will usually reveal whether the prompt, model, data, or route failed.

This is also where SEO and GEO discipline belongs. Google's helpful content guidance favors original value, expertise, and enough information for readers to accomplish their goal. AI answers and search snippets reward the same thing in practice: specific answers, clear entities, strong sourcing, and pages that do not force the reader to keep searching.

For this topic, the useful answer is not a list of model names. Model names age fast. The durable answer is an operating standard: route by task, evaluate with workflow evidence, and keep the human owner close enough to catch drift before customers do.

That is why the fix is not "use Claude for strategy" or "use GPT for speed" as a permanent rule. The fix is a routing habit. If a new model changes the quality curve, you test it. If it beats the current lane on the workflow's actual metric, you move the lane. If it only looks impressive in a demo, it stays out of production.

Related reading: the risk is bigger when teams confuse automation with agency. See multi-LLM routing, reasoning model overhype, and the AI ROI proof problem.

FAQ

What is the jagged AI problem in marketing?+

The jagged AI problem is the pattern where a stronger model improves some marketing tasks while making others slower, more expensive, less on-brand, or less effective. The model is not uniformly better across every workflow.

Why can a better AI model make marketing worse?+

A better model can make marketing worse when it is used for the wrong job. Strategy may benefit from deeper reasoning, but subject lines, lead response, tagging, and short copy often need speed, restraint, and consistency more than maximum intelligence.

What is AI model routing?+

AI model routing is the operating layer that sends each task to the model, prompt, tool, or rule best suited for that workflow. A routing layer considers quality, latency, cost, evidence needs, and brand voice before the request is processed.

How should marketing teams evaluate AI models?+

Marketing teams should evaluate AI models by workflow outcome, not by general benchmarks. The right tests include conversion quality, response speed, source fidelity, brand fit, editing time, and cost per usable output.

Should every marketing team use multiple AI models?+

Not always. Small teams can start with one model, but they still need routing rules. Some tasks may use the same model with different settings, prompts, tools, or review standards. The point is not model sprawl. The point is task fit.

The uncomfortable truth

Your team may not need a smarter model. It may need a smarter traffic system. The model is only the engine. Routing is how the work gets anywhere useful.