I Trained an AI on 3 Years of Ad Data and What It Found Shocked Us

I Trained an AI on 3 Years of Ad Data and What It Found Shocked Us

I Trained an AI on 3 Years of Ad Data and What It Found Shocked Us

The Question That Started It All

For three years, I worked at a mid-sized digital marketing agency. We ran campaigns for e-commerce brands, SaaS startups, and a few household names. Every quarter, someone would ask: "What actually drives conversions?" We'd pull dashboards, slice the data, present a slide deck. The answers were always the same—better targeting, stronger creative, optimized funnels.


But I kept feeling like we were answering a different question. Not what works, but why it works. Not the surface pattern, but the underlying logic.


So I did something a little unconventional. I took three years of our ad performance data—over 1.2 million campaign records, 47,000 individual ads, spanning 212 client accounts—and fed it into a custom training pipeline. Not a generic model. I built a transformer-based architecture tuned specifically for sequential ad performance: impressions, clicks, conversions, spend, creative metadata, audience segments, time-of-day, device, and channel.


I gave the model no labels for "success." No ground truth. Just raw performance signals, and I asked it to find structure.


What it found was... unexpected.

What We Expected vs. What We Found

Before diving into the results, let's be honest about our priors. In digital marketing, the conventional wisdom is:

  • Creative quality is the #1 driver of performance

  • Audience targeting precision correlates with conversion rate

  • Budget scale allows for algorithmic learning and optimization

  • Frequency has a sweet spot, beyond which fatigue sets in

These are not wrong. They're just incomplete.


The model found a pattern that challenged all four assumptions simultaneously.

The Non-Linear Creative Effect

Here's the first finding that surprised the team.


We assumed creative quality had a roughly linear relationship with performance. Better creative → better CTR → better conversions. The data told a different story.


The model identified a bimodal distribution in creative performance. There were two distinct clusters of ads that performed well:

  1. High-production, polished, brand-aligned ads (the kind we'd traditionally invest in)

  2. Low-production, slightly imperfect, almost "amateur" ads

And a large middle band of "competent" creative that performed worse than both clusters.


Let's make this concrete:

Performance by Creative Production Level:

High-Production  ████████████████████  (9.2% CVR)
Medium-Production ████████████  (5.1% CVR)
Low-Production  ████████████████  (7.8% CVR)

The "amateur" cluster—ads with slightly off timing, imperfect copy, even a minor visual inconsistency—outperformed the polished middle. The model traced this to perceived authenticity. Audiences, especially on social channels, had developed a subconscious filter for "this looks like an ad." Slightly imperfect creative reduced that filter. It felt like a recommendation, not an interruption.


This wasn't a targeting effect. The same audience segments saw all three types. The difference was in how the creative felt, not who saw it.

The Targeting Paradox

Second finding: narrower targeting did not always mean better performance.


We'd spent years building hyper-specific audience segments. The model found that performance peaked at a medium level of audience precision, not maximum.

Conversion Rate by Audience Segment Size:

<1M audience  ███████████  (4.8% CVR)
1M-5M audience  ███████████████  (6.3% CVR)
5M-20M audience  ███████████████████  (7.1% CVR)
20M-50M audience  ████████████████  (6.5% CVR)
>50M audience  ██████████  (4.2% CVR)

The sweet spot was around 5-20M reach. Narrow segments underperformed because they were too homogeneous—everyone in the segment had similar needs, similar fatigue patterns, similar ad exposure history. The algorithm couldn't differentiate. Broad segments underperformed because of noise. The middle was where the model found the richest signal-to-noise ratio.


This contradicted the industry narrative that "more precise = better." It suggested that optimal precision depends on the product, the channel, and the audience's own ad exposure history.

The Frequency Curve Wasn't a Curve

Third: we assumed ad frequency had a classic inverted-U shape. Too few impressions → no recognition. Too many → fatigue. Peak in the middle.


The model found something more complex. Frequency effects were conditional on creative type and channel.


For high-production creative on display channels, the classic inverted-U held:

Display / High-Production:
Freq 1-2:  ████████
Freq 3-5:  ████████████  (peak)
Freq 6-10:  █████████
Freq 11-20:  ██████

But for low-production creative on social channels, the peak shifted later and the fatigue set in slower:

Social / Low-Production:
Freq 1-2:  ██████
Freq 3-5:  █████████
Freq 6-10:  ████████████  (peak)
Freq 11-20:  ███████████
Freq 21-30:  ████████

The same "amateur" creative that outperformed polished creative also tolerated higher frequency before fatigue. The model interpreted this as: if the ad feels like content, audiences tolerate (and even prefer) seeing it more. If it feels like an ad, repetition becomes intrusive faster.

The Budget Non-Linearity

Fourth finding: budget scale had a threshold effect, not a gradient effect.


Below a certain spend level, performance was noisy and unstable. Above it, the algorithmic systems (the platform's own optimization) could actually learn and optimize. But there was a knee in the curve.

CVR by Monthly Spend:

$5K:  ████████  (3.1%)
$20K:  ██████████  (4.7%)
$50K:  █████████████  (6.2%)
$100K:  ██████████████  (6.8%)
$250K:  ███████████████  (6.9%)
$500K:  ███████████████  (6.8%)
$1M:  ██████████████  (6.7%)

Performance improved from $5K to $100K. Beyond $100K, it plateaued. And then—this was the part that really surprised us—beyond $250K, it slightly declined.


The model suggested this was because at very high spend, the campaign was no longer optimizing for conversion. The platform's algorithm was being pulled toward impression volume and reach, and the audience composition was shifting toward lower-intent users to fill inventory. The marginal dollar was buying a different kind of audience.

The Hidden Correlation

But the finding that genuinely shocked the team was the fifth one.


The model found a strong correlation between creative update frequency and campaign longevity.


Not just "new creative performs better" (which we knew). But: campaigns that introduced at least one new creative variant every 10-14 days had significantly higher 90-day cumulative conversions than campaigns that ran a static set of 5-10 creatives for the full period.

90-Day Cumulative Conversions:

Static creative (5-10 ads, no swaps):
  ████████████  (1,240 conv)

Light refresh (1 new ad / 30 days):
  ███████████████  (1,580 conv)

Moderate refresh (1 new ad / 14 days):
  ███████████████████  (2,010 conv)

Heavy refresh (1 new ad / 7 days):
  ███████████████████  (1,950 conv)

Extreme refresh (3+ new ads / 7 days):
  ████████████████  (1,620 conv)

The peak was at moderate refresh. Too little refresh → audience fatigue, ad blindness. Too much → the algorithm couldn't stabilize learning, and audiences got confused by constantly changing messages.


The model's interpretation: audiences need novelty to re-engage, but consistency to build brand recall. The optimal strategy is a steady drumbeat of small creative evolution, not a static set or a constant overhaul.

What This Means Practically

None of these findings are revolutionary in isolation. Each one is a refinement of what we already believed. But together, they paint a picture that's meaningfully different from the conventional playbook.


The conventional playbook says:

Invest in high-production creative. Target precisely. Scale budget. Find your frequency sweet spot. Run it and optimize.

The data suggests a more nuanced strategy:

Invest in a portfolio of creative—some polished, some deliberately imperfect. Target at medium precision, not maximum. Scale budget to the knee, not the ceiling. Let frequency be conditional on creative type. Refresh creative on a moderate, steady cadence.

It's less about finding the one optimal setting and more about understanding that optimal is a function of context.

A Note on the Model

To be transparent: this was a single model, trained on a single agency's data. 212 clients, one industry mix, one time period. The findings are suggestive, not proven. A rigorous study would need replication across agencies, industries, and time periods.


But the pattern was consistent enough that it wasn't noise. The bimodal creative effect, the targeting paradox, the conditional frequency curve, the budget knee, the refresh cadence—each one held across multiple client verticals and channels.

The Bigger Picture

What struck me most wasn't any single finding. It was the quality of the insight.


We were looking for "what works." The model was finding "what works, and why, and under what conditions."


That's a fundamentally different question. And it's one that dashboards and A/B tests struggle to answer. A/B tests tell you that Variant A beats Variant B. They don't tell you that the relationship between creative quality and performance is non-linear, that it depends on channel, that it interacts with audience composition, that it has a threshold effect.


A model trained on the full corpus of performance data can find those interactions. Not because it's smarter than a human analyst. But because it can hold thousands of variables in mind simultaneously, and it doesn't confirm-bias toward the story we already believe.


We believed that better creative always wins. The data said: sometimes, slightly imperfect creative wins. And that's a harder lesson to internalize than a clean linear graph.

Final Thought

Three years of data. 1.2 million records. One model. Five findings that challenged our assumptions.


The best part? None of them were findable by looking at any single campaign, any single client, any single quarter. They only emerged when you looked at the whole picture, across time, across channels, across audiences.


That's what data can do that intuition can't. Not replace judgment. But expand it.


And that's worth more than any single campaign optimization.